Watermarks Degrade Clinical LLM Reasoning, Multi-Model Study Finds
TL;DR
- Researchers tested 5 watermarking schemes across 11 LLMs and 7 vision-language models on unimodal and multimodal clinical reasoning tasks.
- Reported failure modes include lexical corruption, hallucinated medical terminology, and misattribution or omission of image findings.
- Authors argue standard aggregate benchmarks can obscure these failures and call domain-specific evaluation a prerequisite for safe clinical deployment.
Adding watermarks to language-model output can quietly damage clinical reasoning without moving the accuracy benchmarks that usually gate deployment, according to a new preprint on arXiv.
The authors benchmarked "5 watermarking schemes across 11 LLMs and 7 VLMs on various tasks spanning unimodal and multimodal clinical reasoning." Watermarking, they write, "can induce substantial degradation across multiple failure modes, including lexical corruption, hallucinated terminology, and amplified misattribution or omission of image findings."
The problem is that standard benchmarks miss this. "The absence of domain-specific analyses, combined with aggregate metrics that miss failures inherent to clinical text, can systematically obscure practical watermark-induced degradations," the paper reports.
To catch what the aggregates hide, the team built a "human-expert-validated pipeline for systematically auditing medical reasoning quality, terminological precision, and induced hallucinations."
The abstract does not break down which of the 5 schemes fared worst, or publish per-model results across the 11 LLMs and 7 VLMs. What it does state, without hedge, is the conclusion: domain-specific evaluation is "a prerequisite for the safe deployment of watermarked models in medicine, where current benchmarks can otherwise mask clinically consequential failures."
Originally reported by paper
Read the original article →Original headline: LLM Watermarks Fabricate Medical Terms, Double Flawed Reasoning—5 Methods, 11 LLMs Tested