By Jay Krivanek, AI Engineering Manager
Anthropic announced that they were watermarking text in order to comply with EU rules. If you have not heard, the EU AI Act's Transparency Code "requires AI companies to mark AI-generated or edited content in a way other systems can identify them”.
Anthropic has chosen the route of watermarking text. A lot of energy has been spent detecting AI. To be clear I am not saying that there is no benefit to watermarking. Certainly being able to tell that a student used AI to write their report can help. I’ve also had my own issues with individuals thoughtlessly throwing a wall of text at me without considering the audience.
Watermarking Solves for the Wrong Problem
However, tackling this challenge with this technique is not only a waste of time, it will generate more waste that an industry, already being heavily criticized for wasting resources, will have to contend with. While Anthropic hasn’t released details on how they will watermark text, one can only imagine it will be done with subtle word nudges. Done in a mathematical pattern so that text remains coherent but a program could decode the pattern to certify that an article is watermarked.
If you’ve spent any time working with AI, you’ve become accustomed to AI-isms. Frequent use of the em-dash, predictable speech patterns, quirky tone as well as the overly verbose nature of AI generated text. This requirement will only compound this problem because now AI nudges words as it writes text. I predict a degradation in quality of output will be felt over time. Anthropic promises that it is imperceptible. So we spend a lot of energy generating text, then spend more energy nudging its words around as it writes.
The Limitations Are Already in the Fine Print
Now you’ve cracked it, we can universally detect AI, right? Not so fast, even Anthropic’s own support document states there are limitations to this watermarking feature. Within their support docs, “The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.” Also Google’s own Nature paper addresses this: “another limitation of generative watermarks is their vulnerability to stealing, spoofing and scrubbing attacks, which is an area of ongoing research.” Not to mention the false confidence that may play out when someone runs content through an “AI Watermark Detector” and it comes up negative. As Anthropic says, “Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed.”
The Watermark Is One Rewrite Away From Gone
Naively, a user could just rewrite the content and the watermark goes away. That same support document says Anthropic is working on allowing outside parties to detect AI using their watermark. What I predict is that there will be services spun up to not only detect but spoil the watermark. Even more simply, a user could run the text through an open weight model with no such watermarking to remove all signs of the watermark. Now we are burning energy to undo the watermark. Even more casually, you could just run a synonym algorithm of your own over the work automatically and say “bye” to the watermark.
The very people you intend to target with this watermark (cheaters) are the very people who will face few consequences under this new detection scheme. Instead your casual Claude users will have to contend with reading watermarked text that only hurts the experience as the war to watermark rages behind the scenes. We don’t even know if this could hurt generated code performance because, so far, we do not know how and what portion of the content is impacted by the watermark algorithm.
AI Literacy Is the Harder, Better Answer
The best case scenario is that it is truly imperceptible, but that promise doesn’t assuage concerns. Even if that pans out, this still doesn’t solve the problem it was intended to solve. What would be more helpful is more AI literacy. How we work with AI, how we properly apply AI. Appropriate use of AI to meet your goals while taking ownership of the generated work. Disclosure of use of AI and concrete policy on AI is the true pathway forward.
