New benchmark says top AI models can only recover 3 to 15 percent of research ideas from bibliographies alone

Started by Poppy51, Yesterday at 10:59 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: New benchmark says top AI models can only recover 3 to 15 percent of research ideas from bibliographies alone   Views(Read 17 times)
Active members in this topic:
Poppy51(1)

Poppy51

A new scientific reasoning benchmark called Reconstruction, published this month, strips a research paper down to just its bibliography and asks frontier language models to guess what idea the paper actually proposed. The results were rough across the board, with individual models landing somewhere between three and fifteen percent accuracy depending on the model and the field being tested.

The benchmark designers tried a more elaborate setup too, running a multi agent Swiss tournament style pipeline where multiple model outputs compete and get filtered down. Even that approach only reached forty two percent, which is a big jump over any single model but still far from anything you'd call reliable hypothesis generation.

What stands out is the narrow spread between models of very different capability levels. If genuine abductive reasoning, meaning forming a prediction from incomplete evidence, were actually happening here, scores should track capability more closely. Instead everything clusters in a tight low band, which suggests whatever mechanism produces a fifteen percent match rate is close to its ceiling already rather than something that just needs a bigger model to fix.

There's a practical angle buried in here as well. A growing industry of AI scientist tools markets language models as engines for generating novel research hypotheses, and a lot of the evaluations behind those claims give the model access to full paper text, author information, or signals that only exist after publication. Stripping all of that away and testing on bibliographies alone is a much harder and arguably more honest test of whether the model is reasoning toward an idea or just pattern matching against things it has already seen


Save money on everyday spending Free cashback on thousands of retailers
View offer