Do AI Text Watermarks Survive a Rewrite?
If you draft with an AI assistant and then rewrite the result before sending it, you may have noticed something new this year: the tools you draft with now quietly mark their own output. Anthropic confirmed on August 11, 2026 that new Claude models watermark generated text by default, worldwide, not only in Europe. Google’s Gemini has watermarked text since 2024. Neither company hides that the practice is happening. What they say much less about is what survives once a person actually edits the draft.
That question matters more for correspondence than for the legal-labeling debate this blog has already covered. A watermark is not a compliance checkbox; it is a statistical signal baked into word choice at the moment of generation. Whether it is still there after you rewrite a paragraph is an empirical question, and the research on it gives a clearer answer than the marketing pages do.
Watermarking Text Went From Experiment To Default
Anthropic’s rollout followed the EU AI Act’s Article 50 transparency obligations, which took effect on August 2, 2026 and require generative AI providers to make their output technically detectable as machine-generated. Anthropic said the watermarking applies across Claude, the Claude API, Claude Code, and related surfaces, and confirmed it uses the SynthID Text approach Google DeepMind published in 2024, the same underlying method already live in Gemini. A detection API is rolling out first to regulators, researchers, and media organizations rather than the general public.
That is a meaningful shift from where things stood even a few months ago, when watermarking was one experimental option among several unreliable detection methods. It is worth reading alongside AI writing detection in 2026: where things actually stand, which covers the broader detector landscape this fits into, and content credentials for text explained, which covers the separate legal question of who has to label what.
What The Watermark Actually Detects
A SynthID-style watermark is not a visible mark, a hidden character, or metadata attached to a file. It works at generation time: when the model has several roughly equally good tokens to choose from, it is nudged toward a particular subset of them in a pattern only the provider’s detector can recognize statistically. Read one sentence and you would never notice. Run enough of the text through the matching detector and a probability score comes back.
Two consequences follow directly from that mechanism. First, the signal exists only in text the model itself produced, token by token. Second, anything that changes enough of those tokens, in a way not dictated by the watermarking bias, weakens the statistical signal the detector is looking for.
What The Research Says Survives An Edit
This is where the actual data is more specific than “editing removes it.”
Google’s own documentation on SynthID acknowledges the watermark is robust to modest changes such as cropping a passage, swapping a handful of words, or light paraphrasing, but that confidence drops substantially under heavier rewriting or translation into another language. Researchers at ETH Zurich’s SRI Lab, probing SynthID-Text directly, found the same pattern from the other direction: text that has been substantially reworded is much harder to attribute back to the watermark with confidence.
A 2026 robustness study on SynthID went further, testing deliberate removal techniques rather than ordinary editing. Its authors found that combining a few specific attack strategies, tournament sampling and larger generation context among them, pushed successful watermark removal above 90% in their tests, and that “stealing” the watermarking pattern first made even simple attacks far more effective. That figure describes an adversary deliberately trying to scrub a watermark, not a person editing a draft for tone, so it is not a direct answer to “what happens when I rewrite this email.” But it establishes the boundary: these watermarks are a probabilistic signal about unmodified generation, not a tamper-proof seal, and they were never designed to survive substantial rewriting.
What This Means If You Rewrite Every AI Draft Anyway
None of this is a reason to think about evading anything. A watermark on an unedited AI draft you never sent is not a problem to route around; it is accurately describing text a model wrote. The practical point is narrower: if your actual workflow is to get a rough AI first pass and then genuinely rewrite it, that workflow is not the same act, statistically or otherwise, as sending the model’s output unchanged.
Before:
Per our conversation, I am reaching out to confirm the details regarding the upcoming project timeline and to ensure alignment on next steps moving forward.
After:
Following up on our call: can you confirm the new deadline by Friday? I want to make sure the team is planning around the right date.
The first version keeps the model’s phrasing almost token for token, which is exactly the shape a watermark detector is built to recognize. The second replaces the structure and the word choices with what you would actually say, which is also, incidentally, the more useful email. A rewrite done for the ordinary reason, because the first draft did not sound like you, changes enough of the text that the statistical basis for a watermark match weakens on its own.
A Wrivio Context for turning a rough AI draft into something you would send could say:
Rewrite this in my own voice: direct, conversational, no corporate filler. Restructure sentences rather than swapping individual words. Keep every name, date, figure, and commitment exactly as written. Do not add sentiment or requests that are not in the original.
Press Ctrl+Shift+Space, paste the draft, and check the diff. The more of the sentence structure actually changes, not just a synonym here and there, the closer the result is to something you wrote rather than something you forwarded.
Common Questions
Does Wrivio remove AI watermarks from text?
Wrivio does not target watermarks at all; it rewrites text for tone, clarity, and register the way any editing pass would. Whether a watermark’s statistical signal survives is a side effect of how much the wording actually changes, not a feature Wrivio implements or claims.
Can a watermark tell someone I used AI to help write an email?
Only if the recipient or their organization has access to the specific provider’s detector, and only for text close enough to the original generation. Detection APIs are currently limited to regulators, researchers, and media organizations rather than being publicly available.
Does a light edit, like fixing a typo, remove the watermark?
Generally no. Both Google’s documentation and independent research describe the watermark as reasonably robust to small changes like word swaps or trimming a passage, while heavier rewriting or translation degrades it much more.
Do local AI models watermark their output?
No. Watermarking is applied by the hosted provider’s generation code, so open-weights models you run yourself on your own hardware are not watermarked, since nothing enforces it once you control the software.
Is watermarking the same as the EU AI Act’s labeling requirement?
They are related but not identical. Watermarking is the technical mechanism providers use to make output machine-detectable; the labeling requirement in Article 50 of the EU AI Act is a separate legal obligation that applies to a much narrower set of published, public-facing content.
Download Wrivio for Windows to turn a rough AI draft into something in your own words, with a diff that shows exactly how much changed.
Read Next
The UK Is Consulting on Workplace Monitoring Rules, and Email Counts
A UK consultation open until 30 September 2026 treats email monitoring the same as algorithmic scheduling. What the proposed scope means for disclosing new tools.
Why Your First Draft Should Not Touch a Public Chatbot
First drafts are the least filtered thing you write, which makes them the riskiest to paste into a public AI tool. Why the rough version leaks the most.
Is It Safe to Use AI for HR and Personnel Writing?
Performance reviews, warnings, and terminations contain some of the most sensitive data you handle. What is safe to run through AI, and what must stay local.
Why AI Search Rewards Exact Numbers and Dates
Vague writing reads fine to humans and badly to answer engines. Why specific figures and dated claims get quoted, and how to add them without inventing them.
This article is filed underPrivacy & Compliance, which has 85 articles.