arstechnica.com web signal

Google SynthID Watermarks Hold Up in Tests, Labeling Still Slips

google ai detection ai-business

TL;DR

  • Google says SynthID has now labeled 100 billion images and videos, plus 60,000 years of audio.
  • Ordinary edits and noise barely dent the watermark, but running an image back through another AI model breaks the signal.
  • OpenAI said in May 2026 it will embed SynthID in all ChatGPT-generated images, joining Nvidia and other adopters.

Watermarking is one of the few tools the industry has agreed to try for telling AI-generated images apart from real ones, and Google's SynthID is the version with the most adoption. Ars Technica put it through a hands-on test, and the finding is that the signal itself is genuinely durable inside Google's own generation pipeline. The problem is not the watermark; it is everything that happens outside the pipeline.

The mechanics are the part that works. A SynthID image watermark hides in the pixels themselves rather than in the file's metadata, which is why the mark survives cropping, added filters, color changes, resizing, lossy compression, and even a screenshot. Ordinary edits and random noise barely dent it. What breaks the signal is anything that rebuilds the pixels from scratch, such as running the picture back through another AI model.

That is the crack that the labeling ambition falls through. Pushmeet Kohli, a Google DeepMind scientist, told Ars that "a technology like this will always be attacked," and independent testing has shown the limits in practice: researchers ran six detection tools against a Nano Banana Pro image and found that simple edits using standard software were enough to drop detection rates to almost zero. DeepMind's own phrasing is that SynthID is not foolproof against extreme image manipulation.

Scale is the other honest caveat. Google says SynthID has now been used to label 100 billion images and videos, plus 60,000 years of audio, and OpenAI announced in May that it will embed the same watermark in all images generated by ChatGPT, joining Nvidia and several other companies. That is meaningful coverage of the biggest generation stacks. It is also nowhere near total, and the reporting does not give you a firm estimate of how much of everyday AI content actually carries a detectable mark, or how the detector behaves on lightly retouched human photos.

The forward-looking read is that watermarking is still worth doing, because a partial signal is useful for the honest parties, platforms, provenance tools, and newsrooms, that want to verify what they can. It is not the deepfake fix it sometimes gets pitched as, and treating the absence of a SynthID hit as proof of authenticity is the mistake to avoid.

Shared on Bluesky by 1 AI expert

  • Jeremy Hsu @jeremyhsu.bsky.social amplified

    @arstechnica.com

    While labeling AI content is a good policy for tech firms, it probably won’t save the rest of us who have to deal with its outputs.

    View on Bluesky →