Subscribe
Technology

Anthropic watermarks Claude outputs but bypass takes hours

Anthropic's watermarking announcement is a compliance move dressed as a trust signal. On 14 August 2026 the company said it would embed invisible watermarks into every output from its Claude models, covering both text and generated files.

5 min read
Anthropic chief executive Dario Amodei speaking on a conference stage
Anthropic chief executive Dario Amodei. Claude's outputs now carry invisible watermarks | Digitally illustrated image
Zara Kincaid
By Zara Kincaid · 2026-08-26

TLDR

Anthropic is watermarking everything Claude produces, embedding an invisible signal in its text and signed manifests in its files to meet European transparency rules. Researchers showed paraphrasing strips the text signal completely, leaving buyers of AI-assisted content with a tool that confirms AI involvement but cannot rule it out.

KEY TAKEAWAYS

01Anthropic announced all Claude models will carry invisible watermarks covering text and generated files.
02Text watermarks use a statistical signal; image files carry cryptographically signed C2PA provenance metadata.
03Older Claude models must comply with EU AI Act Article 50(2) marking rules by 2 December 2026.
04Paraphrasing removes KGW and Unigram text watermarks in 100 per cent of tested cases, researchers found.
05A detected watermark confirms AI generation; its absence cannot confirm the content is human-written.

Useful for confirmation, useless for clearance

Anthropic's watermarking announcement is a compliance move dressed as a trust signal. On 14 August 2026 the company said it would embed invisible watermarks into every output from its Claude models, covering both text and generated files.[1] For businesses buying AI-assisted content at scale, understanding what the tool actually proves matters more than the headline.

Two watermark types, one asymmetric guarantee

When Claude generates text, it weaves a statistical signal into word selection without changing meaning, quality or readability; when it produces supported file types such as .svg.png or .jpg, it attaches cryptographically signed provenance metadata following the C2PA open standard.[2] Anthropic put it plainly: "Watermarking doesn't change the meaning or experience for the person reading it, but if you wanted to check after the fact whether the text was likely generated by Claude, the watermark allows you to do so."[1]

The asymmetry is baked in by design. Anthropic's own documentation states that finding a supported mark indicates the content may have been processed by Claude, but that the absence of a detectable mark does not mean the content was not AI-generated or processed.[2] Detection confirms; absence proves nothing.

The EU AI Act deadline driving the rollout

The announcement traces directly to Article 50(2) of the EU AI Act, which requires providers of generative AI systems to carry machine-readable marks on outputs indicating AI generation.[3] New Claude models launched on or after 2 August 2026 support watermarking at launch. Models released before that date have until 2 December 2026 to comply.[2]

That grace period matters for operators running older Claude versions through APIs. Any workflow that ingests Claude output and resells it as a service carries a compliance clock: four months to instrument detection or surface provenance data to downstream clients.

The bypass problem

The robustness question answered itself quickly. Researchers Saifur Rahman Tamim and Amir Labib Khan found that meaning-preserving paraphrase achieves 100 per cent conditional watermark removal for KGW and Unigram text watermarking methods, and 98.3 per cent removal for SynthID-Text, across 846 valid paraphrase runs covering 15 diverse prompts per method.[4]

Tamim and Khan said: "Out of 846 valid paraphrase runs across 15 diverse prompts per method, every single initially-detected KGW and Unigram text lost its watermark after paraphrasing, 100% conditional removal."[4] For file-based outputs, re-saving an image strips C2PA metadata, a well-documented behaviour in the standard's own literature. The signal survives only if no one touches the file.

What operators should do now

Watermark detection works as a positive-confirmation audit tool when you control the pipeline end to end and the content has not passed through a third-party editor or paraphrasing layer. Run it outside those conditions and the result tells you little.

Publishers and content marketplaces accepting third-party submissions should treat the watermark as one check among several, not a substitute for editorial review. A clean detection result on unmodified Claude output confirms provenance; anything touched since generation is inconclusive. Anthropic's December 2026 compliance deadline for legacy models gives operators a fixed date to audit which Claude versions sit in their stack and whether detection infrastructure needs to be added before then.[2]

FREQUENTLY ASKED QUESTIONS

What types of Claude output carry watermarks?
Text outputs carry a statistical signal woven into word selection. Supported file types including .svg.png and .jpg carry C2PA cryptographically signed provenance metadata attached at the point of generation.
When must older Claude models comply with the watermarking requirement?
Claude models released before 2 August 2026 must comply with the EU AI Act Article 50(2) marking obligation by 2 December 2026.
Can paraphrasing remove the text watermark?
Yes. Research published in July 2026 found that meaning-preserving paraphrase removed KGW and Unigram watermarks in 100 per cent of tested cases and removed the SynthID-Text watermark in 98.3 per cent of cases.
Does the absence of a watermark mean content is human-written?
No. Anthropic's own documentation states that a lack of a detectable mark does not confirm the content was not AI-generated or processed. The tool confirms AI involvement when a mark is found; it cannot clear content as human-written when no mark is found.
Zara Kincaid

Zara Kincaid

Zara Kincaid writes about artificial intelligence and search. Her focus is what happens to businesses when the front page of the internet stops being a list of links and starts being an answer.

Related topics
What's your reaction?

Make us a preferred source on Google

Tap once and our reporting shows at the top of your Google search results and AI answers. You can change this at any time.

Add as a preferred source on Google
Subscribe — it's free