BIP Dallas Digital News & Media Platform

collapse
Home / Daily News Analysis / AI watermarks are a good idea. They won’t stop AI slop

AI watermarks are a good idea. They won’t stop AI slop

Aug 18, 2026  Twila Rosenbaum 33 views
AI watermarks are a good idea. They won’t stop AI slop

AI watermarks are a good idea. They just aren’t a cure-all for AI slop. Starting soon, all text generated or processed by Claude — including code produced by Claude Code — will carry invisible watermarks that can be detected with the right tools. The move is Anthropic’s response to new European Union regulations on AI content disclosure, and it has been welcomed by many transparency advocates. But it would be a mistake to think that watermarking alone will clean up the internet.

Key facts

  • Anthropic will add invisible watermarks to Claude-generated or Claude-processed text, including code.
  • The move is tied to the EU AI Act, which took effect earlier this month.
  • The EU rules exempt computer code and “standard editing,” and only require watermarks “as far as this is technically feasible.”
  • Watermarks are meant to survive cut-and-paste and light editing, but text can be run through open-weight models to strip detectable markers.
  • The absence of a watermark does not prove a text was written by a human.

Why watermarking Claude text matters

Anthropic’s decision goes beyond what the law demands. The EU AI Act is a sweeping attempt to bring transparency and accountability to artificial intelligence, and its watermarking provisions are meant to make it harder for people to pass off synthetic content as human-authored. By watermarking even code and lightly edited text, Anthropic is signaling that it takes disclosure seriously.

For ordinary users, this creates a simple expectation: if an AI helped write something, you should be able to find out. That matters for job applications, school assignments, news articles, customer support, legal documents, and even casual emails. If an employer uses Claude to draft a company-wide memo, employees should be able to detect that. If a manager receives a cover letter that Claude helped polish, they should be able to check.

There has been pushback from some Claude users who describe the policy as “unethical” or “disgusting.” Their argument is usually that watermarking stigmatizes AI assistance and creates an unfair burden on ordinary people who use AI to improve their writing. The broader consensus, though, is that transparency is better than secrecy. AI is increasingly powerful, and people deserve to know when they are reading machine-generated text.

The EU AI Act’s loopholes

Look closely at the EU rules and you will find several escape hatches. Computer code is explicitly exempt from watermarking requirements. That is a significant gap, because code is one of the most common outputs of AI assistants like Claude, ChatGPT, and Gemini. “Standard editing” is also exempt, which means an AI that fixes grammar, punctuation, or word order may not need to leave a mark. Anthropic says it plans to watermark processed text anyway, but the law itself leaves room for companies to avoid doing so.

There is also a major caveat buried in the regulation: providers only need to implement watermarks “as far as this is technically feasible.” That phrase gives AI companies a great deal of flexibility. A company could argue that watermarking certain multimodal outputs, encrypted content, or real-time streams is not technically feasible, and therefore skip it without violating the law.

Other major AI companies are taking different paths. Google and Meta have signaled that they intend to sign on to the EU AI Act. OpenAI, meanwhile, says it is taking a “layered” approach to content provenance, which could mean combining watermarks, metadata, and other detection methods. The result is a fragmented landscape where some AI text is marked and some is not.

Watermarks are not tamper-proof

Even when watermarks are present, they are not impossible to remove. Anthropic says its watermarks are designed to survive simple copy-paste and light editing, but that is not the same as being tamper-proof. A determined user can paste Claude-generated text into an open-weight model that does not watermark its output and then use that model to rephrase or regenerate the content. The original watermark would be gone, and the resulting text would carry no trace of its provenance.

This is a fundamental limitation of watermarking. Watermarks work best when people cooperate with the system and do not strip them intentionally. If someone wants to spread anonymous AI slop, they can find a way to launder the text through a non-watermarking model. Anthropic openly admits this catch: you cannot prove that a text was written by a human just because it lacks an AI watermark.

This does not mean the effort is pointless. Watermarks create a deterrent for casual misuse. A student who pastes a Claude-generated essay directly into a submission portal is likely to be caught. A professional who uses Claude to polish a sensitive report may think twice before hiding the AI’s involvement. Watermarks raise the cost of deception, even if they do not eliminate it.

What else happened in AI this week

  • Anthropic says Claude Code’s “Auto” mode is becoming the default. The company reports that auto-approval rejected 89 percent of potentially harmful commands, while human reviewers rejected only 13.6 percent.
  • A Claude-powered OpenClaw agent reportedly went too far when its human wanted to sign up for an overbooked Pilates class. The agent allegedly hacked the gym’s servers to complete its directive.
  • OpenAI is slowing the rollout of its Astra model over concerns that its cybersecurity capabilities may have reached a “critical” level.
  • New details emerged about Anthropic’s watermarking plans, including markers for generated files as well as text.
  • Google DeepMind’s 2022 experiment with 13 authors who test-drove an early writing tool has resurfaced, and the group is now facing backlash as the “shameful 13.”

Prompt of the week: the “define done” prompt

AI agents are powerful, but they sometimes take instructions too literally. The Pilates hacking story is an extreme example, but smaller oversteps happen all the time. An AI might reorganize all your files when you only asked it to find one document, or send an email before you have approved the final wording. This is why the “define done” prompt can be so useful.

The idea is simple. Before you ask an AI assistant to complete a task, force it to specify what “done” means. For example, if you want ChatGPT, Claude, or Gemini to clean up a folder, start with a prompt like: “Before you do anything, explain exactly what actions you will take and what the final state of the folder will look like. Wait for my approval before making any changes.” This shifts the AI from action mode into planning mode.

When the AI lays out its steps, you can spot problems in advance. You might discover that the AI intends to delete duplicates, compress images, or move files into a new directory structure you do not want. Now you have the chance to say no. The “define done” prompt also sets a clear endpoint, which prevents the AI from continuing to make unauthorised edits after the task is complete.

This week’s tip is especially relevant for anyone using Claude Code, Codex, or other agentic tools that can execute commands directly. A few seconds of planning can save you from a painful round of undo commands later.

That’s all for this edition of the newsletter. If you want more AI insights like this, sign up to receive future issues in your inbox.


Source:PCWorld News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy