An AI spam filter for science just flagged a quarter million cancer papers as potentially fake

Started by Bright Hermit, Jul 18, 2026, 09:40 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: An AI spam filter for science just flagged a quarter million cancer papers as potentially fake   Views(Read 88 times)
Active members in this topic:
Bright Hermit(1) Sam(1) VacantTundra(1)

Bright Hermit

Researchers at Queensland University of Technology have built a machine learning tool that screened 2.6 million cancer research papers published between 1999 and 2024 and flagged more than 250,000 of them, nearly 10 percent, as showing writing patterns closely resembling papers already linked to so called paper mills, businesses that manufacture and sell fake or low quality scientific manuscripts, sometimes built around fabricated data

Led by biostatistics professor Adrian Barnett, the team trained a language model called BERT to recognize the subtle textual fingerprints that repeatedly show up in known paper mill products, recycled boilerplate phrasing, awkward templated sentence structures, and other patterns that don't naturally occur in genuine, individually written research. Tested against verified examples, the tool correctly identified suspected paper mill output roughly 91 percent of the time, and the team is describing it plainly as a scientific spam filter, the same basic concept as an email system flagging suspicious messages based on recognizable patterns

The trend line is what genuinely alarmed the researchers. Flagged papers rose from around 1 percent of the literature in the early 2000s to more than 16 percent by 2022, appearing across thousands of journals published by major publishers rather than being confined to obscure, low prestige outlets. Gastric, bone, liver, esophageal and ovarian cancer research showed some of the highest flagging rates, and the problem appeared in journals with high impact factors just as much as lower ranked ones

Crucially, being flagged is not the same as being proven fraudulent, every single paper the tool identifies still needs expert human review before any conclusion can be drawn, and researchers are explicit that the tool is designed to raise questions rather than answer them. Three scientific journals are already piloting the technology directly inside their editorial process. Barnett framed the stakes plainly, cancer research directly influences clinical trials, drug development and patient care, and if fabricated studies quietly make their way into the evidence base, they risk misleading genuine researchers and slowing real progress for actual patients

Sam

Going from 1 percent flagged in the early 2000s to over 16 percent by 2022 is a genuinely alarming trend line, that's not noise, that's a real and rapidly growing problem in the literature
Posted from my main account

VacantTundra

The careful distinction between flagged and proven fraudulent is the most important nuance here, this tool raises red flags for human experts to investigate, it doesn't hand down a verdict on its own

Save money on everyday spending Free cashback on thousands of retailers
View offer