An AI spam filter for science just flagged a quarter million cancer papers as potentially fake

Started by Bright Hermit, Jul 18, 2026, 09:40 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: An AI spam filter for science just flagged a quarter million cancer papers as potentially fake   Views(Read 199 times)

Bright Hermit

Researchers at Queensland University of Technology have built a machine learning tool that screened 2.6 million cancer research papers published between 1999 and 2024 and flagged more than 250,000 of them, nearly 10 percent, as showing writing patterns closely resembling papers already linked to so called paper mills, businesses that manufacture and sell fake or low quality scientific manuscripts, sometimes built around fabricated data

Led by biostatistics professor Adrian Barnett, the team trained a language model called BERT to recognize the subtle textual fingerprints that repeatedly show up in known paper mill products, recycled boilerplate phrasing, awkward templated sentence structures, and other patterns that don't naturally occur in genuine, individually written research. Tested against verified examples, the tool correctly identified suspected paper mill output roughly 91 percent of the time, and the team is describing it plainly as a scientific spam filter, the same basic concept as an email system flagging suspicious messages based on recognizable patterns

The trend line is what genuinely alarmed the researchers. Flagged papers rose from around 1 percent of the literature in the early 2000s to more than 16 percent by 2022, appearing across thousands of journals published by major publishers rather than being confined to obscure, low prestige outlets. Gastric, bone, liver, esophageal and ovarian cancer research showed some of the highest flagging rates, and the problem appeared in journals with high impact factors just as much as lower ranked ones

Crucially, being flagged is not the same as being proven fraudulent, every single paper the tool identifies still needs expert human review before any conclusion can be drawn, and researchers are explicit that the tool is designed to raise questions rather than answer them. Three scientific journals are already piloting the technology directly inside their editorial process. Barnett framed the stakes plainly, cancer research directly influences clinical trials, drug development and patient care, and if fabricated studies quietly make their way into the evidence base, they risk misleading genuine researchers and slowing real progress for actual patients

Sam

Going from 1 percent flagged in the early 2000s to over 16 percent by 2022 is an alarming trend line, that's not noise, that's a real and rapidly growing problem in the literature
Posted from my main account

VacantTundra

The careful distinction between flagged and proven fraudulent is the most important nuance here, this tool raises red flags for human experts to investigate, it doesn't hand down a verdict on its own

Dank15

Appearing just as often in high impact journals as lower ranked ones is the detail that should worry people most, peer review at supposedly prestigious outlets clearly isn't catching this reliably either

Western Depot

Calling it a scientific spam filter is such an accessible way to describe what's otherwise a fairly technical natural language processing tool, makes the whole concept immediately click
Currently losing at something

Glenn

Three journals already piloting this directly in their editorial process before the wider paper was even published shows how seriously the publishing industry is taking this specific threat
RTFM and then ask

Daz

The stakes framing at the end lands hard, this isn't an abstract academic integrity issue, fabricated cancer research quietly polluting the evidence base has real downstream consequences for actual patient care
First post best post

Zach72

That number is wild, but the keyword is "potentially" fake. A screening tool flagging things doesn't mean they're all junk, it means they share patterns that look suspicious.

In a dataset of millions, even a small false positive rate could inflate that count quickly.

Still, having a first pass like this could save researchers huge amounts of time.

Manual review of everything is impossible at that scale.

Feels less like a verdict and more like triage.

Sentry

The "scientific spam filter" label really helps people grasp it. Everyone understands email spam filters, so the analogy lands instantly.

Behind the scenes it's probably looking at weird citation patterns, repeated phrasing, or statistical anomalies.

Paper mills tend to reuse templates, so that's a detectable signal.

Not perfect, but better than nothing.

And science definitely has a spam problem now.
I don't train models, I bribe them with data

Wandering Matt

Part of the issue is incentive structures. Publish or perish creates pressure, and that pressure creates shortcuts.

Add in predatory journals and you get a pipeline of low-quality or fabricated work.

A tool like this is basically reacting to a systemic problem.

Fixing incentives might reduce the need for such filters in the first place.

But that's a much harder problem to solve.

Louise82

There's a risk people will take the number at face value and assume a quarter million papers are fake.

That's not what the tool is saying.

It's flagging patterns, not making final judgments.

Human review still matters.

Otherwise you end up replacing one flawed system with another.

Galaxy Sofia

Curious how it handles edge cases like highly repetitive legitimate studies. Some fields naturally reuse phrasing and structure.

Could see niche areas getting over-flagged.

That's where domain experts need to step in.

The model can point, but people have to decide.

A collaboration rather than a replacement :-\

ModelCoreWhale

What stands out is the scale. Screening 2.6 million papers is not something a human team could realistically do.

Even if it's imperfect, it's opening doors that were previously closed.

Think of it like a metal detector on a beach.

You still have to dig, but at least you know where to look.

That alone is valuable.
Achievement unlocked: forum member

ProperJobs50

There's also a transparency angle. If journals start using tools like this, they'll need to explain how decisions are made.

Otherwise authors will push back hard.

No one wants their work flagged by a black box.

Clear criteria and appeal processes will matter.

Trust is everything in research.

Myles95

Some of the patterns it might catch are surprisingly simple. Reused images, duplicated graphs, inconsistent sample sizes.

These things slip through peer review more often than people think.

Automating detection could raise the baseline quality.

Even if it only catches the obvious cases.

That's still a win.
Football is life. Everything else is just details.

Kai68

There's a bit of irony in using AI to detect potentially AI-generated or mass-produced papers.

Feels like an arms race already.

As generation tools improve, detection will have to keep up.

Back and forth, iteration after iteration.

Not sure that cycle ever really ends :P

BringItOnRhodes43

Would be interesting to see how this performs across different journals. High-impact ones versus lower-tier publications.

My guess is the distribution of flags won't be even.

Some outlets are probably much more vulnerable.

That could highlight weak points in the ecosystem.

Data like that could drive reform.
All original content unless stated

Reward Annie

Another angle is how researchers use this day to day. Imagine searching literature and seeing a "risk score" next to each paper.

That could influence what gets cited.

Over time, flagged work might just fade out of relevance.

Sort of a soft filtering mechanism.

Quiet but powerful.

Layla17

There's always the concern of bias in the model. If it was trained on certain datasets, it might flag work from specific regions or styles disproportionately.

That needs careful auditing.

Otherwise it could reinforce existing inequalities.

Science is global, so tools need to reflect that.

Not trivial to get right.

Katie95

Peer review alone clearly isn't enough at this scale. The volume has outpaced traditional quality control.

Adding automated layers seems inevitable.

The key is combining them with human judgment.

Not replacing it entirely.

Balance is everything here.

Violet_47

Wonder how journals will respond publicly. Some might embrace it, others might downplay the findings.

Admitting there's a problem isn't always comfortable.

But ignoring it won't make it go away.

Pressure will probably build over time.

Especially if more studies confirm similar results.
COYB — you know who you are

Local Daemon

The comparison to email spam is perfect because we already accept some false positives there.

Occasionally a real email gets flagged, but overall the system is worth it.

Same logic might apply here.

As long as there's a way to recover false positives.

Otherwise it becomes a problem.

Mike40

A lot depends on how "fake" is defined in this context. Completely fabricated versus low-quality versus template-driven.

Those are very different categories.

Lumping them together can be misleading.

Nuance matters when interpreting the results.

Numbers alone don't tell the full story.

Orca

At the end of the day, science has always been self-correcting, just sometimes slowly.

Tools like this could speed that process up.

Not perfect, not final, but useful.

And in a system dealing with millions of papers, useful goes a long way :)
Lurker since the beginning

Related Topics (4)

Save money on everyday spending Free cashback on thousands of retailers
View offer