Anthropic starts marking Claude-generated text and files with invisible watermarks and C2PA metadata
Anthropic has published details of a marking system built into Claude that embeds an invisible watermark in generated text and attaches signed provenance metadata to generated image files, according to a Claude Help Center article, "How Claude marks AI-generated content."
What's new
The system uses two separate techniques depending on content type. For text, Anthropic says: "When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself." The watermark is embedded at generation time, inside the model's token selection process, rather than appended afterward, and Anthropic notes it "will travel with the text when it's copied and pasted elsewhere, and may persist through some editing."
For image files, Claude attaches a different kind of marker. Per the article, "When Claude generates a supported file type, such as a .svg, .png, or .jpg, it will attach signed provenance metadata...which follows the Coalition for Content Provenance and Authenticity (C2PA) open standard" — the same open standard backed by Adobe, Microsoft, and camera makers for tracking image origin and edit history.
Coverage is tied to model release date: "Claude models launched on or after August 2, 2026 will support machine-readable marking at launch," with Anthropic adding that it is "also working to add marking support to Claude models released before that date."
Context
Anthropic is explicit that the system is not a detection tool. The article cautions that finding a Claude mark only indicates content may have been processed by Claude, not that a human wrote none of it — marks can fail to survive heavy editing, paraphrasing, translation, mixing with other writing, or very short passages, and older models that predate the feature won't carry a mark at all.
The move lands alongside a broader industry push toward content provenance standards as AI-generated text and images become harder to distinguish from human work at scale. C2PA in particular has become a common reference point: image and audio generators from other labs have adopted C2PA-style provenance metadata or the related SynthID watermarking approach as a lighter-weight alternative to outright content detection, which remains unreliable for text.
Why it matters
By building marking directly into Claude's generation process rather than offering a bolt-on detector, Anthropic is betting on provenance over detection as the more durable answer to AI-generated content questions — the mark travels with the content instead of requiring a separate tool to classify it after the fact. That distinction matters for institutions worried about AI-generated misinformation, academic dishonesty, or synthetic media: a watermark that survives copy-paste gives platforms and downstream tools a hook to check without needing Anthropic's exact model weights. But Anthropic's own limitations list — no protection against paraphrasing, translation, or older models — means the system narrows the problem rather than solving it, and works best as one signal among several rather than a definitive answer to "did an AI write this."
Corroborating sources
- Support.claude
https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
“When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself.”