A student writes an essay without using artificial intelligence, submits it for evaluation and receives an accusation of using ChatGPT. The reason may have little to do with whether the student used AI and more to do with how predictable their writing appears to a detection tool.Research led by Stanford’s James Zou found that AI detectors could incorrectly flag a large proportion of essays written by non-native English speakers. In the study described in the source material, more than 61% of 91 essays written by students preparing for an English proficiency examination were classified as AI-generated, on average, by seven popular detectors.The findings raise questions about the use of AI detection scores in schools and universities, particularly when students’ academic records and credibility could be affected by an automated assessment.
Why AI detectors may mistake simple English for ChatGPT writing
The Stanford team examined essays written by students preparing for an English examination and compared the results with writing by American eighth-grade students.According to the account of the research, 89 of the 91 essays by non-native English speakers were flagged by at least one of the seven detectors. By contrast, the American students’ essays received predominantly human classifications.The explanation lies in how many AI detectors assess writing. They use patterns such as predictability and the likelihood of particular words appearing in a sentence to estimate whether a text was generated by a machine.AI models often select words that are statistically likely to follow the preceding text. Students writing in a second language may also rely on familiar vocabulary and straightforward sentence structures because they are more comfortable with them.This can create an unexpected problem: writing that is simple, grammatically predictable and entirely human can resemble the patterns detectors associate with AI-generated text.For students, this means that the way they express themselves, rather than the actual process through which they wrote an essay, could influence how a detector classifies their work.
ChatGPT rewrote the essays, and the false flags fell
The researchers also explored whether changing the vocabulary and structure of the flagged essays would affect the results.When ChatGPT was asked to rewrite the essays using more sophisticated vocabulary, the reported detection rate fell from approximately 61% to around 12%.The result highlights a fundamental difficulty with AI detection. If rewriting a human essay with an AI tool can make it less likely to be flagged, a detector’s score cannot reliably establish who actually wrote the original text.The problem is not limited to third-party tools. OpenAI launched its own AI text classifier in January 2023, but discontinued it six months later because of its low accuracy. The company had acknowledged that the tool could incorrectly classify human-written text and fail to identify AI-generated content.
An award-winning writer was accused of using AI
The consequences of false detection can extend beyond the classroom.In May, Jamir Nazir, a 62-year-old writer from Trinidad, won a regional title in the Commonwealth Short Story Prize. Soon after, an AI detector called Pangram reportedly classified his story as entirely AI-generated.The accusation circulated online, prompting people to question the authenticity of his work.However, Nazir has neuropathy and reportedly dictates his stories into his Android phone, editing them in small sections. The prize organisers subsequently examined his drafts, character notes and earlier versions of the story, including an ending he had discarded.After a month-long review, they cleared him and named him the overall winner.The case illustrates why a detection score can be an inadequate substitute for examining how a piece of writing was created. The original drafts and working notes provided evidence that a percentage generated by software could not.
What should schools and universities do instead?
For teachers and academic institutions, the challenge is to distinguish between evidence of AI use and a detector’s prediction.The article’s source material notes that Turnitin advises educators against using its AI detection score as the sole basis for taking action against a student.A more detailed review can involve examining drafts, revision histories, research notes and the student’s ability to explain the arguments and writing choices in their submission.Such checks can provide context that an automated score cannot capture, particularly when students have different levels of English proficiency or use accessibility tools to produce written work.As AI becomes more common in education, the distinction between detecting a pattern and establishing academic misconduct becomes increasingly important.For students, the central concern is straightforward: a detector can raise a question about an essay, but its score alone does not establish who wrote it. Before an accusation affects a student’s academic record, the evidence behind the work deserves to be examined.Disclaimer: This article is based on the research and cases described in the supplied source material. TOI Education has not independently verified the study’s findings or the individual allegations and accounts.