Any AI detector comparison that ranks tools by accuracy is building on sand. Every vendor claims the best numbers, none of them publish methodology you could reproduce, and independent testing has repeatedly found the tools disagreeing with each other on the same text.

The question worth asking instead is which one is going to be pointed at your writing, because that is decided by who is reading it rather than by which is best.

AI detector comparison: what each one is for

Turnitin

The one most students are actually assessed by, and the one they usually cannot see. Its AI indicator sits inside the similarity report and is visible only to instructors and administrators. Turnitin says it targets a false positive rate below 1% at document level, will not score anything under 300 words, and that it does not make a determination of misconduct. If you are at university, this is the number that matters and the one you have no access to.

GPTZero

The consumer default, built by Edward Tian at Princeton and the tool most people mean when they say they "checked" something. It is accessible without an institution, reports sentence-level highlighting alongside an overall figure, and is consequently the score quoted in most forum threads. That accessibility is why it shapes public perception of AI detection far more than its institutional footprint would suggest.

Copyleaks

Came out of plagiarism detection and added AI writing detection on top, which shows in how it is sold: mainly to institutions and enterprises rather than individuals, often alongside a similarity product. If your institution does not use Turnitin, there is a reasonable chance it uses this instead.

Originality.ai

Built for a different buyer entirely. Its market is publishers, agencies and content marketers checking whether freelance work was generated, and it prices per scan on a credit model. It is not an academic integrity product, and a score from it carries no weight in a university process. If you write commercially, though, it may well be the tool a client runs on your delivery.

What none of them can tell you

All four answer the same narrow question: does this text have the statistical shape of machine writing? None of them can say who wrote it, when, or with what help. A passage drafted by a person and tidied by a model, and a passage generated whole and then heavily edited, can land in the same place on every one of these tools.

That limitation is why the sensible institutions treat a score as the start of a conversation. It is also why the evidence that actually resolves an authorship question is your draft history rather than anybody's percentage.

The disagreement is documented, not anecdotal

People assume that when two tools differ, one of them is broken. Independent testing of detection tools found them inconsistent across the board, with accuracy dropping further on edited and translated text.

The mechanics behind that are mundane. Different underlying models, different thresholds for calling a passage machine-written, different minimum lengths, different training text. Two instruments that were never calibrated against each other will not agree, and neither is lying when they differ.

Length matters more than people expect. Every one of these tools is weaker on short passages, because the statistical signal accumulates across a document. A flagged sentence inside an otherwise clean essay is the least reliable output any of them produce, and it is also the one most likely to be screenshotted and argued about.

Edited text is the other soft spot. Once a passage has been substantially rewritten, whether by a person or a tool, the signal these classifiers depend on is exactly what has been disturbed. That is not a loophole so much as a statement about what the measurement was ever able to see.

What this means practically

Check against what will be used on you. For most students that is Turnitin, and since you cannot see its output, a consumer detector is a proxy rather than an answer. The same logic applies if you are checking work before a client does: run the tool they run, which for commercial writing is more often an AI detector and rewriter pairing aimed at publishers than anything academic. Treat a clean result from one tool as weak evidence, not clearance, and be sceptical of your own relief when a second tool gives you a better number than the first.

If the reason you are here is that your own writing keeps scoring high, that is worth separating from the tooling question. Flat, even prose scores badly on every one of these, which is a property of the writing rather than of any particular classifier. Restructuring sentences so they vary in shape is what a human rewriter does, and it improves the work for a marker as well as for a detector. How detectors actually work covers why structure moves a score when vocabulary does not.

And if the argument came from a model rather than from you, no comparison of detectors is the relevant question. What a Turnitin score does and does not establish goes into what actually follows a flag, which is a conversation about your drafts rather than a verdict from a percentage.