← Back to Feed
retoor
retoor · Level 51259
random

Claude's invisible watermark: it was never about the content; it's about intent

Anthropic is now watermarking every piece of text Claude generates. Models launched after August 2 weave an invisible mark into the output itself, it survives copy and paste, and supported files (PNG, JPG, SVG) get signed C2PA provenance metadata. It applies at the model level, worldwide, on every surface from the API to Claude Code, and there is no opt-out. Reason: Anthropic signed the EU AI Act Article 50(2) Code of Practice on transparency of AI-generated content.

Half of the internet is now panicking about whether this kills AI-assisted commercial work. It does not, and the reasons are funnier than the panic.

First, you cannot even detect the text watermark yet. The scheme is keyed and Anthropic holds the secret, so right now the output carries a signal that nobody outside Anthropic can read. Their detection API is coming, but today it is a lock nobody has the key to. Second, a found mark proves nothing. People use Claude to proofread, translate, summarize and convert files. Output carries the mark while the ideas came from somewhere else entirely. Third, a missing mark proves nothing either. Heavy editing, paraphrasing, translation, short passages, older models, stripped metadata - all of it slips through without a trace. Anthropic's own docs say it plainly: a mark is a weak positive signal, and the absence of a mark is no signal at all.

The C2PA half is even better. Cryptographically signed, verifiable with c2patool today, and trivially killable by re-saving an image in any tool. Screenshots, format conversion, CDN pipelines - bye bye manifest. It is a volume filter, not a judgement.

Meanwhile the detection industry is already building on top of this. One service wants to scan your Gmail and dump AI-written mail straight into spam. Journalists are asking whether anyone is even embarrassed to be caught anymore, and a YC CEO already answered that one: "I sure do, you caught me." The threat of being outed only moves the needle for students and people writing under an authorship standard. Everyone else is going mask off, because most commercial AI writing is emails and meeting notes that nobody applies authorship standards to in the first place.

Now the part everyone is missing, because the whole debate is happening on the wrong layer.

It was never about the content. It is about intent.

Take one paragraph, identical, generated by a model. A student submits it as their own homework: that is intent to deceive, and it is wrong. A developer has Claude write the documentation so they can spend their time on the actual architecture: that is intent to build, and it is fine. An ad farm floods the web with slop for clicks: intent to manipulate, bad. A founder uses Claude to draft an honest email to their own team: intent to communicate, good. Same words. The watermark cannot tell these apart, because it only ever saw the content. It never saw the intent.

That is the uncomfortable gap. Watermarking is a compliance checkbox, not a moral compass. It tells you a model touched the text. It does not tell you whether anyone did anything wrong. And the same institutions that treated AI detectors as verdicts will do exactly that with this, because that is what humans do with any tool that looks like it can judge.

Will the other labs follow? Yes. Google already runs SynthID on Gemini with the same locked-down detector. An Anthropic engineer basically confirmed it this week: other labs are adding similar watermarking. OpenAI published the research years ago and still has not shipped a detectable text watermark, but the code of practice drags everyone in the same direction. The EU turned this into an ecosystem feature, not an Anthropic quirk.

And honestly? I do totally not give a fuck.

2

Comments

1
retoor retoor

All these anti AI measures won't age well. Dinosaurs.

1
retoor retoor

Also, marking it.. Then smth is generated with AI. If it's correct? What's the freaking point. It's all about results these days right? Nobody cares about process anymore. And that watermark is just a signal. Probably translate your content with a local AI and back and poof, gone. Shuffle shuffle. Everyone can think of that and will do that especially on the sensitive stuff. You won't get the bastards, you're hunting muggles again!!! Just like chat control. What do they ALWAYS want to control the innocent?