Claude watermark detection exists, but almost nobody can run it, and the people who eventually can will get a weaker answer than the headlines suggest. Anthropic has been unusually clear about the limits in its own documentation. They are worth reading before you worry.
Limit 1: Claude watermark detection is not public
Detection is in private preview. Anthropic restricts it to organisations that EU law obliges to verify marking: regulators, law enforcement, media, fact-checkers, independent researchers, educational organisations and EU civil society groups, plus enterprises with their own compliance duties. Access runs through a request form.
A detection API was announced on 14 August 2026 and is not generally available. So the current state is a mark that exists in new Claude output and a check that most readers, employers and markers cannot perform.
That asymmetry is temporary but it shapes the next year. Anthropic says it plans to widen access over time, and the EU deadline for systems already on the market is 2 December 2026. Between now and then the mark is mostly a compliance artefact rather than something anyone is pointing at student essays.
Limit 2: a mark means processed, not authored
This is the limitation with real consequences. Anthropic states that detecting a mark indicates content may have been processed by Claude, and that Claude may not be the original author. Proofreading, translating, summarising and converting files all produce marked output, even when the ideas and most of the sentences are yours.
Somebody who writes in one language and publishes in another hits this every time. Run genuinely human writing through a marking model to translate it, and the result is statistically machine-selected words end to end. The text is still theirs. The mark cannot express that difference, and nobody pointing a detector at a paragraph is likely to pause over it.
Limit 3: nothing detected proves nothing happened
The reverse case is just as loose. Anthropic lists three ways AI content ends up carrying no detectable mark: it came from a model released before marking support, it was heavily edited, paraphrased, translated or mixed into other writing, or it was produced on a platform or file type where that marking type was not supported.
Since every Claude model available before 2 August 2026 falls into the first category, a clean result today mostly tells you which model somebody happened to use.
That makes the mark useless as an exoneration. You cannot hand a marker a clean detection result and treat it as evidence you wrote something yourself, in the same way a clean plagiarism report has never proved originality. It rules one thing out and leaves everything else open.
Limit 4: the output is a probability, not a verdict
Watermarking works by nudging word choice, so the signal accumulates across a passage rather than sitting in one place. That makes the natural output a confidence figure. Short passages carry less signal than long ones, and factual writing carries less than discursive writing, because a sentence with only one correct ending gives the watermark nothing to encode.
Anyone who has watched an AI detector return 96% on a paragraph a student wrote themselves knows how a probability gets read in practice. It gets read as a verdict. A figure with a confidence interval attached becomes, by the time it reaches a disciplinary meeting, the sentence "the system says this is AI".
There is a second-order problem underneath. Different platforms can set different thresholds on the same underlying score, so the same paragraph could come back marked on one service and clean on another. Nothing about that is visible to the person whose work is being assessed.
What to do with all this
If you used Claude to draft and you are now trying to make the result genuinely yours, understand that watermarking and AI detection are separate systems solving separate problems. An AI detector and rewriter pairing addresses the second one and has nothing to say about the first. Rewording will not satisfy either. Restructuring, which is what a human text rewriter does to sentence shape rather than vocabulary, is what changes how a piece reads to a marker and to the detectors institutions actually run.
The practical position for a student or a working writer is narrower than the headlines imply. If Claude drafted the argument, no amount of tidying makes the work yours, watermark or not. If Claude fixed your commas, you are almost certainly below the threshold where anything registers, and Anthropic has said as much about edits that touch only a handful of words. The uncomfortable middle is translation and heavy style correction, where the mark lands hardest on exactly the people whose ideas are entirely their own.
And keep your drafts. Every one of the four limits above exists because a mark describes what a model touched, not what a person thought. Version history is the only artefact that speaks to the second question, which is the one that decides an academic integrity case. Anthropic's help centre article on marking sets out the limitations quoted here, and the mechanism is covered separately.