Pakkā

The Problem With Watermarking AI

Article 50 of the EU AI Act is now in effect. Among its provisions, it requires providers of generative AI systems to make AI-generated or manipulated content detectable via machine-readable marking. The European Commission’s Code of Practice translates this into concrete guidance: for free-form text exceeding 200 tokens, the preferred compliance method is an imperceptible watermark.

The underlying goal is reasonable. Generative AI has made it far harder to determine the origin of digital content. A photo, video, audio recording, or article may now come from a machine rather than a person. Offering a reliable way to establish provenance is valuable, particularly when the material is intended to inform or shape public opinion.

Yet the EU is committing a category error.

Knowing that content was produced by AI is not the same as knowing whether it is trustworthy. This distinction is critical because technical provenance is becoming easier to establish at the same time that genuine authenticity is growing harder to assess.

A watermark can indicate that a model likely produced a given paragraph. It cannot reveal whether that paragraph is factually correct. It cannot show whether the writer relied on AI for the entire argument or only for polishing a sentence. It cannot confirm whether the publisher endorses the content. And it cannot tell anyone whether the material deserves belief.

In short, the watermark answers the wrong question.

The deeper difficulty is that text differs fundamentally from images. With an image, provenance can be an intrinsic property of the file: we can record that a photograph was taken by a specific camera, edited in Photoshop, and later altered by an AI tool. That history carries useful information.

Text works differently because writing is inherently iterative and transformative. Someone can prompt a model for a full essay, revise it by hand, translate it, run it through another model, reorder sections, delete large portions, and publish the result under their own name. At what stage does the text cease to count as “AI-generated”?

Any legal rule must draw a boundary somewhere. Technology must then produce a signal that survives beyond that boundary. This is where watermarking becomes problematic.

The system is no longer simply generating text; it is quietly modifying the output so that a third party can later detect the machine’s involvement. The change may be statistically minute and leave quality unaffected. Still, the text now carries information intended for someone other than the user—whether or not the user consents to its presence.

This is an odd paradigm for the future of software. More importantly, it is unlikely to fix the real problem.

Anyone determined to hide AI assistance can simply paraphrase the output, route it through another model, translate it, edit it extensively, or switch to a system outside the regulated ecosystem. As a result, watermarking is most likely to catch ordinary users—the very people who had little interest in deception to begin with.

This exposes the broader mistake. Regulators are attempting to revive an outdated internet assumption: if content appears human, treat it as human-created. That assumption is vanishing, and no watermark can restore it.

A more constructive approach is not to embed secret identifiers in every piece of generated text. It is to create an environment in which provenance is available when it matters, while trust is established by other means. That requires open standards, verifiable origin data, editorial accountability, clear authorship, and reputation systems. It means applying something closer to zero-trust principles to digital information: never assume authenticity simply because material looks authentic.

Article 50 correctly identifies a genuine challenge. Watermarking, however, should be treated as a limited provenance tool rather than a remedy for the authenticity crisis unleashed by generative AI.

The internet will soon contain more synthetic content than ever. The decisive question is not whether every machine-written sentence can be flagged. It is whether we can design systems in which knowing a piece of content’s origin is helpful—without pretending that origin alone tells us what to believe.