New benchmark says top AI models can only recover 3 to 15 percent of research ideas from bibliographies alone

Started by Poppy51, Aug 23, 2026, 10:59 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: New benchmark says top AI models can only recover 3 to 15 percent of research ideas from bibliographies alone   Views(Read 39 times)

Poppy51

A new scientific reasoning benchmark called Reconstruction, published this month, strips a research paper down to just its bibliography and asks frontier language models to guess what idea the paper actually proposed. The results were rough across the board, with individual models landing somewhere between three and fifteen percent accuracy depending on the model and the field being tested.

The benchmark designers tried a more elaborate setup too, running a multi agent Swiss tournament style pipeline where multiple model outputs compete and get filtered down. Even that approach only reached forty two percent, which is a big jump over any single model but still far from anything you'd call reliable hypothesis generation.

What stands out is the narrow spread between models of very different capability levels. If genuine abductive reasoning, meaning forming a prediction from incomplete evidence, were actually happening here, scores should track capability more closely. Instead everything clusters in a tight low band, which suggests whatever mechanism produces a fifteen percent match rate is close to its ceiling already rather than something that just needs a bigger model to fix.

There's a practical angle buried in here as well. A growing industry of AI scientist tools markets language models as engines for generating novel research hypotheses, and a lot of the evaluations behind those claims give the model access to full paper text, author information, or signals that only exist after publication. Stripping all of that away and testing on bibliographies alone is a much harder and arguably more honest test of whether the model is reasoning toward an idea or just pattern matching against things it has already seen


StringTheory55

Forty two percent from a whole tournament of agents versus three to fifteen from single models is actually a pretty big multiplier even if the ceiling is low. Diminishing returns from more compute doesn't mean zero returns

Southern Jay

Does the benchmark control for field at all? I'd expect wildly different scores between something like biomedical research where citation patterns are dense and something more novel or interdisciplinary where the bibliography tells you a lot less

Ruby92

I want to push back a little on calling this a ceiling. A tight low band across models could just as easily mean the task itself is underspecified, since plenty of legitimate research directions could plausibly follow from the same bibliography.

Multiple valid hypotheses fitting one set of citations isn't a flaw in the models, it's just how science works most of the time. Recovering the exact idea the original authors had in mind is a much narrower target than recovering a reasonable idea
Not financial advice. Not medical advice. Just vibes.

HiddenFreddy81

This lines up with something I've noticed using these tools for my own literature review work. They're great at summarizing what's already been said and decent at connecting related papers, but the moment you ask for a genuinely new angle they tend to just recombine existing framings dressed up as something fresh.

That's not necessarily useless, recombination is part of how real research ideas get generated too, but it is a long way from the AI scientist pitch some of these startups are selling to investors. Calling this benchmark a wake up call feels about right

BlackMamba35

Forty two percent with a multi agent tournament also costs a lot more compute and API calls than a single query, worth keeping in mind before anyone gets too impressed by that number

RobVanDam10

I'd be curious whether human researchers do meaningfully better on the same stripped down task. If a domain expert also only gets to fifteen or twenty percent from bibliography alone, the benchmark says less about AI limitations and more about how much information actually lives in a citation list versus the paper's actual text

Related Topics (5)

Save money on everyday spending Free cashback on thousands of retailers
View offer