AI writing detectors keep falsely accusing students, and the harm isn't evenly distributed

Started by BrokenDave72, Yesterday at 04:09 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: AI writing detectors keep falsely accusing students, and the harm isn't evenly distributed   Views(Read 59 times)
Active members in this topic:
BrokenDave72(1)

BrokenDave72

Theres a growing and genuinely well documented backlash against AI writing detectors, the tools schools and universities have leaned on to flag suspected AI generated student work, and the core problem is that these detectors keep flagging real human writing as machine generated, with the harm falling disproportionately on specific groups of students

The scale of the false positive problem is genuinely significant, Businessweek tested leading detectors on 500 human written essays and found 1 to 2 percent were falsely flagged as AI generated, which sounds small until you consider the volume of student work submitted every year, at the University of Kansas specifically, the companys own chief product officer estimated its detector incorrectly flags around 1 percent of overall documents and 4 percent of individual sentences, extrapolated across the universitys student body that works out to an estimated 38,500 students falsely accused of submitting AI written work

The bias in who gets falsely flagged is the part that should concern people most, a Stanford study ran seven widely used detectors over 91 essays written by non native English speakers and found an average false positive rate of 61.3 percent across those detectors, with 97.8 percent of the essays flagged as AI generated by at least one tool, the same detectors read US eighth grade essays with near perfect accuracy, meaning careful, formal, straightforward prose, exactly the style non native speakers and neurodivergent students are often taught to write in, reads as machine made to these systems while more casual native speaker writing sails through, separate research from Common Sense Media found Black students are more than twice as likely as their white peers to be falsely accused of using AI

The real world consequences of these false positives are genuinely severe, students have had graduations delayed, academic records damaged and relationships with teachers permanently strained, one particularly striking secondary effect is what some educators are calling a Cobra Effect, a student who wrote her own essay in her own words started running her writing through AI tools defensively after hearing that stylistic features like em dashes were rumored to trigger AI detectors, the tool designed to prevent AI use became the actual reason she started using AI, precisely backwards from its intended purpose

The institutional response has been genuinely significant too, a growing list of universities including MIT, Yale, Vanderbilt, Northwestern, Johns Hopkins, Georgetown, NYU and Indiana have either banned AI detection tools outright or strongly discourage relying on them as the sole basis for an academic integrity accusation, Vanderbilt specifically calculated that even Turnitins own claimed sub 1 percent false positive rate would mean roughly 750 papers a year get wrongly flagged just at their institution alone, the University of Pittsburghs teaching center put it bluntly, stating it could not endorse the tools given the substantial risk of false positives and the consequential issues such accusations imply, the emerging consensus across academia increasingly seems to be that these detectors should function, at most, as a prompt for a genuine conversation with a student, never as standalone proof of wrongdoing
sudo make me a sandwich

Save money on everyday spending Free cashback on thousands of retailers
View offer