How the detection actually works
Language models choose each word by predicting the most probable next token. Detectors run that logic in reverse: they measure how closely a passage matches what a model would most likely have produced. Text where nearly every word is the expected one reads as machine-generated. Text with more variation—unusual phrasing, uneven rhythm, unexpected specifics—reads as human.
This is why the score is an estimate rather than a finding. It describes a statistical property of the writing, not its origin. A human can write predictably, and a model can be prompted to write unpredictably.
How accurate is it?
Turnitin has publicly stated that the indicator should be treated as a starting point for a conversation rather than a verdict, and that scores on short submissions are less reliable. Independent testing has repeatedly produced false positives on documented human writing, with elevated rates for writers whose first language is not English and for plain, highly structured prose.
The consequences of that have been real. Several universities—Vanderbilt among the most widely reported—turned the feature off, citing accuracy concerns and the risk of wrongly accusing students. Others continue to use it but require corroborating evidence before any accusation proceeds.
Why honest writing sometimes scores high
- Uniform sentence length throughout the document.
- Formal register with no contractions anywhere.
- Repeated sentence openings across paragraphs.
- Standard academic phrasing taught in writing courses.
- Heavy proofreading that removes natural irregularity.
Notice that most of these are things students are explicitly taught to do. Careful, conventional academic writing shares real statistical features with generated text, and no detector can separate the two reliably.
If you have received a high score
Start with evidence of process rather than argument about the number. Version history in Google Docs or Word shows the document being built over time and is far more persuasive than disputing a percentage. Keep your outlines, notes, and sources. Ask to discuss the work: explaining why you structured an argument a particular way is something only the author can do convincingly.
It also helps to know which parts of your draft read as formulaic. A single document-level percentage gives you nothing specific to respond to. Penlify scores each sentence separately, so you can see whether the issue is concentrated in a formulaic introduction or spread across the piece—and understand which habits produced it.
What no tool can do is guarantee a particular Turnitin result. Detectors differ from one another and change with updates, and any product promising a guaranteed score is overselling. The realistic goal is understanding your own writing, not chasing a number.