arxiv.org web signal

LLM watermarks degrade medical reasoning, ETH paper finds

TL;DR

  • ETH Zurich researchers benchmarked 5 watermarking schemes across 11 LLMs and 7 vision-language models on clinical reasoning tasks.
  • On Phi-4-14B with distortionary SynthID, the rate of correct answers backed by flawed reasoning jumps from 11.4% to 26.3%.
  • Fabricated medical entities rise up to +39.2 percentage points, while watermarking only the final answer is 'essentially free.'

Watermarking large language models, the technique regulators and vendors are pushing as a way to trace AI-generated text, can quietly wreck the reasoning underneath when those models are used for medicine, according to a paper from ETH Zurich and Berlin Institute of Health researchers posted on arXiv.

The team, led by Melanie Rieff with Robin Staab, Thibaud Gloaguen, Stefan Hegselmann and Martin Vechev, calls it "the first rigorous study of how LLM watermarks affect medical performance," benchmarking five watermarking schemes across 11 LLMs and 7 vision-language models on unimodal and multimodal clinical reasoning tasks. Their headline result is that accuracy scores can look fine while the reasoning beneath them degrades. On Phi-4-14B under the distortionary variant of SynthID, the paper reports, "the rate of correct answers backed by flawed reasoning more than doubles on several models (e.g. 11.4%→26.3% on Phi-4-14B with distortionary SynthID)."

The failure modes the authors surface are ones a standard benchmark would miss. "Fabricated medical entities increase by up to +39.2 pp" in some configurations, and in the worst cases the paper says up to 40% of a model's clinical output can "rely on mutually exclusive diagnostic statements." On multimodal tasks it describes amplified misattribution and omission of image findings even when aggregate scores hold steady.

For reasoning models the authors draw a sharper line. "Full WM degrades reasoning quality on every axis except the final answer," they write, while restricting the watermark to just the final answer is "essentially free." Three researchers on our Who's Who list shared the preprint within days of upload.

The paper's closing argument treats domain-specific evaluation as a prerequisite before watermarked models are deployed in medicine, warning that "current benchmarks can otherwise mask clinically consequential failures."

Shared on Bluesky by 3 AI experts