An AI writes one word at a time.
At each step, it looks at everything written so far and works out which word should come next. It doesn't get one answer. It gets a whole list, with a score for each one.
Here's the key thing most people miss: very often, several words score almost the same.
"The weather was cold and ___"
"Grey" works. "Overcast" works. "Damp" works. The AI has no strong reason to prefer one. So it picks one at random - essentially a coin flip.
That coin flip is already there. It's why you get slightly different wording every time you run the same prompt.
Every text watermark in this family works by taking control of that coin flip.
The first serious attempt at this, published in 2023, is usually called the green list / red list method.
At each step, a secret key splits the AI's whole vocabulary into two halves: a green list and a red list. Then the system pushes the AI toward green words and away from red ones.
Detection is easy afterwards. Count how many green words appear. Far more than half? It was probably AI.
But look at the cost. The AI is no longer choosing from its own preferences - it's choosing from a list someone else drew up. If the perfect word for the sentence landed on the red list, tough. You get the second-best word instead.
Do that thousands of times and you get exactly what people fear:
Writing that feels slightly off
Lost precision, because the exact right word got blocked
Real factual errors, if the only correct answer was red
Researchers have a word for this: distortionary. It changes what the AI would have said.
This is real. A 2026 study found that watermarked models could hallucinate extra details, comply with requests they should have refused, or refuse harmless ones — and the effects held up even after controlling for how fluent the text looked.
So the fear is well-founded. It's just aimed at the wrong method. Google and Anthropic don't use this one.
Here is where nearly every blog post on this topic goes wrong.
You'll read that modern systems apply a "gentle nudge" to word probabilities - pushing "blue" down from 70% to 45% and "clear" up to 50%.
That isn't what happens. That description is just the green-list method with softer language. Nudging probabilities is the exact thing that causes the damage.
Google's method, called SynthID-Text, does something different. It doesn't touch the AI's scores at all. It changes how the winner is picked.
Think of it as a small audition.
The AI produces its list of options, completely untouched.
Instead of drawing one word, the system draws several candidates from that list - all of them words the AI was genuinely happy to use.
Those candidates are put into a knockout bracket, like a tennis draw.
A secret key decides who wins each match.
The last one standing gets written.
Every word in that bracket was already a word the AI wanted to use. The key doesn't invent options - it only decides which of the good ones comes out.
That's the whole trick. And it's why this version can be non-distortionary: mathematically, the output still follows the AI's own preferences. The green list can't manage that at any setting. The research was published in Nature in October 2024.
Anthropic has confirmed Claude uses a version of this same method.
Google didn't just claim it doesn't. They measured it, on real traffic.
Before rolling it out, they checked quality across nearly 20 million live Gemini responses. No measurable drop in how users rated the answers. Human reviewers comparing watermarked and unwatermarked text side by side couldn't tell the difference either.
Anthropic is equally direct: nothing is added to the text, there are no hidden characters, no extra cost, and no slowdown. The watermark won't push Claude toward a word it wasn't already considering.
This is the question every developer asks, and the honest answer is more interesting than "the system protects you."
The watermark needs a real choice to hide in. Where only one answer is correct, there's nothing to work with.
2 + 2 = has to be 4.
A function name has to be exact.
A closing bracket has to be a closing bracket.
Run the audition in that situation and every single candidate is the same word. Whoever wins, the output is identical. No signal gets embedded - and no damage gets done either.
Notice what didn't happen there. No safety system kicked in. No "entropy threshold" fired. There was simply nothing to choose between.
You'll read that these systems have a special mechanism that detects fragile code and switches the watermark off to protect you. That mechanism doesn't exist. It doesn't need to.
The practical result: your code carries very little watermarking. What's there mostly sits in places where wording is genuinely free - comments, variable naming, prose in documentation. Anthropic says the effect on the code itself is negligible.
Three honest limits.
Short text. Detection is a statistics problem. Each word contributes a tiny piece of evidence, and you need enough pieces to be confident. A one-line reply doesn't have enough. This isn't a rule anyone wrote - it's the same reason you can't judge a coin from three flips.
(You'll see "200 words" or "50 words" quoted as the threshold. Neither number appears in Google's paper, DeepMind's blog, or Anthropic's write-up. Someone made them up and everyone copied it.)
Very factual text. A list of dates, a recipe, a set of measurements. Same reason as code - few genuine choices, so little to mark.
Heavy rewriting. The watermark lives in the specific words. Google says it survives light edits and mild paraphrasing, but confidence drops sharply when text is thoroughly rewritten or run through another language.
One correction worth noting: if you ask Claude to translate a watermarked text, the result is still watermarked - because Claude chose every word in the new version. It's passing it through someone else's translator that breaks the pattern.
Here's what surprised me when the Claude announcement landed. The backlash wasn't about quality at all.
It was about credit.
The watermark marks the words the model chose. So if you write a piece yourself and ask Claude to edit it heavily, some of those words are now the model's. Anthropic acknowledges this directly: light proofreading usually leaves too little to register, but heavier editing may.
That put a particular group of people in an awkward spot - the ones who write their own work and use AI only to clean it up. On Reddit, the strongest argument wasn't technical. It was: I supplied the idea, the context and every round of refinement. The AI was a tool. Why does the tool get the fingerprint?
Reactions ran from "the dumbest thing I've ever heard" to a genuine worry that the mark becomes a scarlet letter on writing people intend to sell.
That's the real debate. Not is the AI dumber now, but what does a watermark actually prove?
It proves: this text probably passed through that company's model at some point.
It does not prove: that the AI wrote it. That a human didn't do the thinking. Who owns it. Or whether using it was appropriate.
That gap is where the arguments will happen for the next few years -in classrooms, newsrooms and HR departments, mostly conducted by people who won't read the paper.
Worth knowing: this isn't a voluntary product decision. It's the EU AI Act, and around 190 signatories agreed to the same code of practice. It's live for new Claude models from August 2, 2026, applied worldwide rather than just in Europe. Expect the rest to follow.
If you've read other explanations of this, three claims are worth throwing out:
"It nudges probabilities toward synonyms." No. It leaves probabilities alone and changes how the winner is selected. Different family of method entirely.
"It detects fragile code and switches off to protect accuracy." No such switch exists. Where there's one right answer there's nothing to mark, automatically.
"You need 50 / 200 words for it to work." Nobody published those numbers. More text means more confidence, on a smooth curve, with no magic line.
Watermarking isn't a silver bullet. Google says so themselves - they call it a building block. But the engineering is genuinely elegant, and it deserves to be described accurately.
Anthropic, Watermarking Claude's text output - https://www.anthropic.com/news/claude-text-watermark
Google DeepMind, Watermarking AI-generated text and video with SynthID - https://deepmind.google/blog/watermarking-ai-generated-text-and-video-with-synthid/
Google DeepMind, SynthID product page - https://deepmind.google/models/synthid/
Dathathri et al., Scalable watermarking for identifying large language model outputs, Nature, October 2024 - the tournament sampling paper
Kirchenbauer et al., A Watermark for Large Language Models, 2023 - the original green list method