Anthropic’s Watermarks: A Regulatory Comfort, Not a Real Solution
The Performative Act of AI Watermarking
Anthropic, a prominent builder of large language models, recently declared its intention to watermark text generated by its Claude models. This move, echoing similar pledges from firms like Suno and commitments by an industry consortium including Google and Microsoft, arrives in direct response to the European Union’s AI Act Transparency Code, which became effective on August 2. On the surface, it appears to be a responsible, proactive step towards content provenance in the age of generative AI. However, this rush to technically tag machine-generated text fundamentally misinterprets the nature of how information spreads and evolves online, creating a comforting but ultimately flimsy shield against the very problems it purports to solve.
The core issue isn’t whether an AI model can embed an imperceptible signal into its output; it’s whether that signal survives contact with human intent and the digital wild. Anthropic notes its watermarks will “travel with the text when it’s copied and pasted elsewhere, and may persist through some editing.” The crucial phrase here is “some editing.” Text is uniquely fluid. Unlike a deepfake video or an AI-generated image, which requires specific tools and expertise to alter without obvious degradation, text can be rephrased, paraphrased, summarized, or simply rewritten by a human or another AI in seconds, effortlessly stripping away any embedded metadata or structural pattern.
The Practical Folly of Digital Provenance
Consider the actual vectors of misinformation. It rarely involves the wholesale copy-pasting of a single, raw AI output. Instead, it’s a process of reinterpretation, slight modification, and iterative sharing across platforms. A headline here, a paragraph there, a series of bullet points — each can be AI-generated, then lightly human-edited, then spread. The incentive for companies like Anthropic to implement these measures now is clear: regulatory appeasement and brand protection amidst growing public distrust. They need to demonstrate compliance and responsibility, even if the proposed technical solution is akin to putting a single, flimsy lock on an open gate.
Substack CEO Chris Best’s concern about “Claudefishing”—the use of AI to generate deceptive content—is real. But a watermark at the model level, as Anthropic suggests, is a forensic tool, not a preventative one for textual manipulation. The C2PA open standard, which Anthropic will use for files, offers robust content credentials, but its application to raw text faces a distinct challenge. Text is designed for modification and reinterpretation. Even a subtle linguistic watermark might be detectable by specific forensic tools, but it offers little practical deterrence against a malicious actor or even an unwitting user who simply edits a few words or runs the text through a second, unwatermarked large language model.
A False Sense of Regulatory Security
This entire approach risks lulling regulators and the public into a false sense of security. The EU AI Act’s Transparency Code, laudable in its intent, demands that AI-generated content be identifiable. But a “technical solution” for text that can be bypassed with minimal effort provides only the illusion of identification. What truly matters is the meaning and intent behind the content, not just its origin marker. The focus on watermarks diverts attention from the more challenging, yet essential, work of fostering media literacy, developing robust platform policies against coordinated inauthentic behavior, and improving human-led content moderation.
While the industry plays a critical role, the burden of proof cannot rest solely on undetectable digital tattoos. The challenge with generative text, unlike images or audio, is its sheer fungibility and the ease with which it can be anonymized, recontextualized, or subtly altered by humans or other AI systems. The sharpest observation here is that watermarking text created by an LLM is like stamping a serial number on individual grains of sand in a desert and hoping it prevents someone from scooping up a handful. It’s a measure that speaks to the technical capabilities of detection rather than the human realities of dissemination. Until we address the human and systemic factors enabling misinformation, these technical fixes, while legally expedient, will remain largely symbolic.