You paste an essay into a detector and it says 96% AI. What did it actually find? Not a database match, not a record of your ChatGPT session. A statistical model looked at your sentences and judged them too predictable to be human. Understanding that judgment is the most useful thing you can know about AI detection, whether you're trying to pass one or disputing a false flag.
Detectors Are Classifiers, Not Plagiarism Checkers
A plagiarism checker like Turnitin's similarity report compares your text against a database of web pages, journal articles, and past submissions, then reports overlap. It can point to a source and say: this sentence appears there.
An AI detector has no such database to compare against, because ChatGPT produces different text for every user. Instead it's a classifier: a model trained on millions of human and AI writing samples that learns the statistical fingerprints separating the two. Scoring your essay just answers one question, does this look more like the AI pile or the human pile? The output is a probability, not a match, which is why it can never prove anything and can flag an essay you wrote entirely yourself.
A detector can never prove anything. It can only measure how predictable your sentences are.
Perplexity and Burstiness, the Two Core Signals
Perplexity measures how predictable each word is. Detectors run your text through a language model and measure how surprised it is by every word choice. AI text has low perplexity almost by definition, since it was produced by a model choosing likely words. A human writing "the results of the study indicate that" scores low; "what the numbers actually show is messier" says the same thing with far less predictability. This is also why synonym-swapping tools fail: an equally likely synonym leaves perplexity untouched, a point we cover in our guide to making ChatGPT text sound human.
Burstiness measures variation in sentence length and structure. Human writing has natural rhythm: a long, winding sentence followed by a short one. AI writing tends toward the middle, moderate length, sentence after sentence, built on similar grammatical frames. A detector doesn't need to understand your argument to measure this; it just checks how flat the distribution is.
What Modern Detectors Add on Top
Perplexity and burstiness were the whole story in 2023. The major commercial tools (Copyleaks, Turnitin's AI module, GPTZero) are now fine-tuned classifiers trained on labelled human and AI text, retrained as new models ship. They pick up subtler patterns too: a recognizable intro-three-points-conclusion shape, uniform paragraph lengths, and stock phrases (delve into, it is important to note, in today's rapidly evolving landscape), cataloged in our list of ChatGPT phrases to avoid. Most also score sentence by sentence, which is why a human draft with AI-polished paragraphs often comes back as a patchwork of flags.
What Actually Triggers a Flag
No single signal flags a document. It's the density and consistency that matters: sustained low perplexity across many sentences, flat burstiness reinforcing it, and classifier confidence above whatever threshold that vendor set, which is why the same essay can score 12% on one tool and 60% on another. Detectors are also unreliable under roughly 150 to 200 words, so most refuse to score short text or attach low confidence to it.
Why Detectors Get It Wrong
False positives. Anyone whose prose is highly regular can trip a detector: non-native English speakers who learned formal sentence templates, scientific writers whose genre demands formulaic structure, and students writing in a careful, plain register. Independent studies have repeatedly measured elevated false-positive rates for non-native writers, and every honest vendor acknowledges it.
False negatives. Text that's been genuinely restructured, new rhythms, varied lengths, reordered logic, loses the statistical profile the classifier was trained on. Structure moves scores; surface edits don't. That asymmetry is the entire basis of effective rewriting, and it's why quick paraphrase tricks fail against GPTZero while structural rewriting works.
Vendors advertise 98 to 99% accuracy, but those numbers come from clean lab conditions: pure AI versus pure human text, long documents, familiar models. Real-world text (mixed authorship, edited drafts, short passages) performs worse, so institutions are told to treat a score as a signal, not evidence.
What Detectors Cannot See
A detector reads final text and nothing else. It can't see your version history, your notes, or the hours you spent, and it can't tell "wrote with AI" from "writes like AI." Everything it knows comes from word statistics in the pasted text, which is why serious academic-integrity processes pair a score with human review or an oral defense rather than acting on the number alone. For the policy side, see what universities actually say about ChatGPT use in 2026.
Using This Knowledge
Restructured sentences, varied rhythm, and unpredictable phrasing move a score. Synonyms, added typos, and reordered paragraphs with identical sentences inside them don't. And the best protection against a false flag is your own drafts and version history, since a probability means nothing against documented process.
A human rewriter tool like AI Rewriter shows a live AI-detection score on every version it produces, so you can watch structural changes move the number in a way word swaps never do. First rewrites are free, no card required.
What Reddit Gets Right, and Wrong
Search any detector question and half the results are Reddit threads, so it's worth checking the popular claims against the mechanics above. "Detectors are snake oil" is half right: the false-positive problem is real, but the underlying signals are measurable, not random. "Add typos or weird spacing" is wrong, since cosmetic damage leaves sentence structure untouched and classifiers are trained on exactly that trick. "Run it through a paraphraser" is outdated; that worked against 2023 detectors, not current ones. And "they can't prove anything anyway" is true about the score and a terrible safety plan, since integrity cases proceed on your whole writing history, not just the number.