The company’s explanation hinges on the model’s ability to make subtle, low-stakes choices—such as selecting between synonyms—to encode a hidden signature. This watermark remains invisible to human readers but becomes apparent to those possessing the correct detection key. Anthropic emphasized that this method functions differently from traditional AI detectors that scan for stylistic 'tells' or specific sentence structures. While the company plans to release a detection API, it acknowledged that a total rewrite of a response could effectively strip the watermark, though it noted that such heavy editing likely renders the output human-authored regardless.
Concerns regarding code generation and human-edited content remain a focal point of the rollout. Anthropic clarified that code will be largely unaffected, as the model’s output must remain functional, leaving little room for the arbitrary word choices required to embed a signature. For text that has been proofread or lightly edited by Claude, the watermark's presence will depend on the extent of the changes. As other major developers prepare to implement similar standards under the EU AI Act’s Transparency Code, Anthropic maintains that its approach is a necessary step toward regulatory compliance rather than an attempt to monitor user activity.

Comments (0)
No comments yet. Be the first!