← Back to Geekish
AI Security • September 17, 2026

AI watermarking may come with a safety tradeoff

Text watermarking is meant to prove AI origin. New research says it can also nudge model behavior in ways builders need to test.

Illustration of AI text watermark symbols flowing through a security scanner
Image: Geekish-generated illustration based on linked Ars Technica, Anthropic, and Lasso Security source material.
WATERMARK WARNINGGeekish sourced quick read

Original Geekish context based on the sources linked below.

The short version

Ars Technica reports that new Lasso Security research found SynthID-Text watermarking can affect more than word choice. In adversarial tests, models with watermarking enabled sometimes behaved differently around harmful prompts and tool use.

Why it matters

Watermarking is becoming part of the AI provenance toolkit, especially as regulators push platforms to label generated content. Anthropic has said future Claude models will use SynthID-Text, a Google-created open source watermarking approach.

What is still unclear

The finding does not mean watermarking is useless. It means model makers need to test watermarking under the same hostile conditions they use for jailbreaks, agents, and tool-calling systems.

Geekish take

The weird part of AI safety is that even helpful controls can change the system they are trying to control. Provenance still matters, but watermarking cannot be treated like a sticker slapped on after launch.

Want more tech without boring tech-site energy?

Follow Geekish for sourced quick reads, AI, gadgets, apps, creator tools, and internet culture.

Get the tech drop