Every post on this site says an AI detector score is a probability rather than proof. That is a claim worth testing instead of repeating, so we tested it against our own tools. Nine AI-written passages, three rewrite modes, two commercial detectors, 72 scored texts, all measured on 19 September 2026. Rewriting did pull the average score down on both detectors. The average turned out to be the least informative number in the whole run.

What did the AI detector reliability test measure?

Nine passages of 159 to 184 words, generated by Claude across five genres: essay introductions, lab report discussions, literature analysis, technical explainers and a personal statement. Each was scored by Originality.ai and ZeroGPT. Each was then rewritten three times through this site's own ai to human rewriter in Default, Simplify and Elevate modes, and every rewrite was scored again by both tools. That is 27 rewrites and 72 detector readings. One ZeroGPT call returned an unreadable response and is recorded as an error rather than quietly folded into an average as a zero.

Did both detectors catch the raw AI text?

One did, one did not. Originality.ai scored all nine untouched passages at exactly 1.00, a clean sweep. ZeroGPT scored six of the eight it could read at 1.00, but returned 0.28 and 0.00 for the two lab report discussions. The second of those deserves a moment: a passage written entirely by a language model, never edited, came back with a 0% AI score. Before any rewriting and before any technique, one of the two detectors called machine text human. Published testing across a wider set of detection tools has reported the same weakness.

Do the two detectors agree with each other?

Often they do not. Across the 27 rewrites the two scores differed by an average of 0.34, and by 0.50 or more on eight of them. The widest gap was 0.96, where a rewritten technical explainer scored 0.00 on Originality.ai and 0.96 on ZeroGPT. A rewritten personal statement ran the other way, 0.93 against 0.00. These are not narrow disagreements about a borderline case. They are two tools reading identical words and returning opposite verdicts.

What did a rewrite that split the detectors look like?

The widest disagreement is worth showing in full. The original technical explainer began: "A database index is a data structure that improves the speed of data retrieval operations at the cost of additional storage and slower writes." The Default rewrite returned: "A database index is an additional data structure created over top of a database table for the purpose of speed up data retrieval from a database table at expense of more storage space usage and slower writes." Originality.ai scored that 0.00. ZeroGPT scored it 0.96. The rewrite is plainly worse English than the sentence it replaced, and the two tools could not have disagreed more completely about it.

Does rewriting actually lower the score?

On average yes, and by a lot. Originality.ai fell from 1.00 to 0.32 in Default mode and 0.24 in Elevate. ZeroGPT fell from 0.79 to 0.39 and then 0.14. Quoted alone, those numbers would support a confident marketing claim. The standard deviations undercut it: roughly 0.40 on a scale that only runs from 0 to 1. The scores did not slide down together. They scattered toward the ends, with most texts landing near 0.00 or near 1.00 and few anywhere in the middle.

Why does the average hide the real finding?

Because the same passage scores differently depending on which mode rewrote it, in no consistent direction. One literature analysis scored 0.02 after Default, 0.98 after Simplify and 0.00 after Elevate. A personal statement went 0.00, then 0.93, then 0.05. In Simplify mode the Originality.ai mean was 0.55 while the median was 0.93, which is the signature of a split distribution rather than a shift. Some texts cleared completely, others did not move at all. One reading on one text tells you about that reading.

What does this mean if you have been flagged?

It means the number records one tool's opinion on one day. If two commercial detectors can differ by 0.96 on identical text, and one can score unedited model output at zero, a score alone cannot carry the weight an academic integrity process tends to place on it. That is the argument for treating a human text rewriter with a live score as a comparison instrument rather than a target. Your drafts, your version history and your ability to explain your own argument stay the durable answer to a flag, as university processes themselves suggest.

What can this test not tell you?

A good deal. The passages came from Claude, not ChatGPT or Gemini, and detector behaviour does not transfer automatically between models. Nine passages is a small sample. The rewrites came from one model at one set of parameters, and that model is not deterministic: a pilot run scored one passage at 0.43 where the full run scored the same passage and mode at 0.87. Every figure here comes from a single run. Read it as a demonstration that these scores are unstable, which it shows clearly, and not as a ranking of either tool.