Introduction
Since language models began producing texts indistinguishable from those of humans, the question of their traceability has become central. On August 11, 2026, Anthropic announced the integration of an "invisible" watermark into Claude's text outputs for models launched from August 2, 2026, with a progressive rollout to older models (TechCrunch). The idea is appealing: thanks to discreet statistical patterns, it would finally be possible to know whether content was generated by an AI.
The promise raises several questions.
Can we trust a watermark that only its creator knows how to detect? Does it resist transformations of the text? And above all, what is the legal value of a signal that remains probabilistic?
How Anthropic's Watermark Works
The mechanism relies on a form of statistical steganography: a pattern is woven directly into the generated text, through the distribution of certain tokens, without relying on separate metadata.
Anthropic explains that "because the watermark is part of the text, it travels with the text when it is copy-pasted elsewhere, and can persist through certain modifications".
The marking applies at the model level, so regardless of the product used (Claude, the API, Claude Code, Claude Cowork, Claude Tag), including via third-party platforms.
For generated files (images in .png, .jpg, .svg formats), Anthropic uses the open C2PA standard, which associates digitally signed provenance metadata — far more credible. Unlike a cryptographic signature verifiable by anyone with a public key, detecting the textual watermark requires a specific tool.
Anthropic indicates it plans to provide detection tools to third parties, but this remains to be made concrete (Business Insider).
The Technical Limits
A statistical watermark is far from an invincible defense. Anthropic acknowledges that textual marking can be erased if the text is heavily rewritten, and that a file's metadata can disappear if its format changes or if a simple screenshot is taken (Forbes).
In other words, the robustness announced against copy-paste guarantees nothing against substantial rewriting, translation, or editorial reworking. Anthropic also notes an important nuance: the mark only signals that Claude "had a hand" in the content, not that a model generated the entirety of the text (Fortune).
Another limit: the policy applies only to models launched from August 2, 2026, with earlier models not yet covered, leaving a gray zone over a significant portion of current usage (Wavect). Finally, as long as the detection tool remains proprietary and not universally accessible, an information asymmetry persists: only the issuing company or its designated (but not yet certified) partners can authenticate the signal with certainty.
The Legal Context: Article 50 of the AI Act
This initiative responds directly to Article 50 of the European AI regulation (AI Act), which entered into application on August 2, 2026. The text requires providers of generative AI systems to ensure their outputs (audio, image, video, or text) are marked in a machine-readable format and detectable as artificially generated (EU AI Act).
Article 50(4) adds a distinct obligation for disseminators: explicitly disclose any deepfake, as well as any AI-generated text published to inform the public on matters of general interest, except where human editorial control applies (Paul Weiss).
The associated Code of Practice specifies that, for text, statistical watermarking remains an interim solution, textual marking technology being intrinsically more difficult to standardize than for image or audio (AI Policy Desk).
A deadline has also been granted until December 2, 2026, for systems already on the market before August 2, and interoperability of detection mechanisms between providers is only expected by February 2, 2027 (Paul Weiss).
Failure to comply with these obligations can lead to substantial sanctions under the AI Act. The Code also requires, in practice, the combination of at least two marking techniques (tamper-proof metadata and imperceptible watermark) and prohibits users from removing this metadata (Paul Weiss).
Legally Fragile Evidence
This is where the crux of the problem jumps out at the discerning technologist: a statistical signature is not direct evidence in the legal sense.
In civil or commercial litigation, evidence must in principle be clear, reproducible, and verifiable by both parties; yet, as long as the detection tool remains proprietary, neither the judge nor the opposing party can contest or corroborate it independently.
The AI Act itself distinguishes two different logics: Article 50(2) imposes marking detectable by technical tools, while Article 50(4) requires disclosure directly perceptible by a human. A law firm also points out that an Article 50-compliant watermark does not, by itself, satisfy the human disclosure obligation of Article 50 (Two Birds).
This distinction confirms that the technical watermark is above all a tool for regulatory compliance and large-scale detection, not an autonomous evidentiary instrument before a court. Moreover, since the burden of proof generally falls on the party alleging a fact, anyone claiming that a given text was generated by Claude would have to rely on additional elements (independent expertise, context, admissions, other material clues) — the mere statistical indicator being contestable in itself, notably as long as no recognized independent certification body exists to audit these methods.
In Short, Still an Announcement Effect...
Anthropic's announcement constitutes a technical advance that allows it to remain compliant with the regulation now in force since August 2, 2026. It demonstrates an effort to meet the AI Act's transparency requirements, with a deliberately global deployment, not limited to the European Union (Euronews).
But on the evidentiary level, this watermark remains one technical indicator among others, not irrefutable proof. It can be erased by substantial rewriting, it does not yet cover all models, and its verification depends on tools whose availability and reliability for independent third parties remain to be demonstrated.
As long as no independent and interoperable certification framework is fully operational (interoperability only being expected in February 2027), this "algorithmic signature" will remain more a tool of internal traceability and compliance than autonomous legal proof opposable before a civil or commercial court.
Sources : TechCrunch, Euronews, Forbes, Fortune, Business Insider, Wavect, EU AI Act, Paul Weiss, AI Policy Desk, Two Birds
