NYU paper: AI scribes mix small talk into 35% of clinical notes
TL;DR
- In 576 patient-clinician dialogues, frontier LLMs inserted small-talk exchanges into 35% of generated clinical notes.
- 3.7% of those frontier-model notes misattributed the asides or used them in clinical reasoning.
- Background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts and contaminated 5.3% of open-weight notes.
In 576 patient-clinician dialogues, frontier LLMs inserted small-talk exchanges into 35% of the clinical notes they generated. In 3.7% of those notes, the models misattributed the asides or used them in clinical reasoning.
The result comes from a new arXiv paper by Krithik Vishwanath, Eric K. Oermann and colleagues at NYU Langone Health's Department of Neurosurgery. "Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning," the authors write. "Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter."
A second experiment ran 57 mock recorded consultations with background speech from a separate patient encounter mixed in at -10 dB. That leaked audio reached 48.2% of transcripts, and contamination carried into 5.3% of downstream notes generated by four open-weight models.
The paper does not stop at measurement. Its authors propose a "dual-encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning." In mechanistic experiments, "restoring clean activations in these heads improved reasoning in distracted contexts, whereas suppressing them impaired reasoning even on clean clinical vignettes." Filter the distraction and you may lose some of the diagnosis.
Mean quality scores shifted by at most 0.20 points on five-point scales, small enough that routine chart review would plausibly miss the contamination at scale.
Originally reported by paper
Read the original article →Original headline: Ambient AI Scribes Contaminate 35% of Clinical Notes With Small Talk, 48% With Adjacent-Patient Audio