arxiv.org web signal

Clinical AI paper: transcripts show visit flow, not patient state

TL;DR

  • The paper analyzed 439 real-world clinical encounter transcripts spanning 134 hours, including 245 ENT transcripts paired with 273 PROM surveys.
  • Conversational phase structure is recoverable from transcripts, but patient state measured via PROM scores for voice, cough, and swallowing is only partially observable.
  • The authors call the pattern an 'observability asymmetry' and warn against transcript-only inference of human state.

A new arXiv paper by Lily Chen, Ted Mau, Michael Gensheimer, Brian Anthony Nuyen, Nancy Jiang, and James Zou puts an empirical limit on what an AI listening in on a doctor visit can actually infer. The team looked at 439 real-world clinical encounter transcripts spanning 134 hours, including 245 ENT transcripts paired with 273 PROM (patient-reported outcome measure) surveys. Annotation ran through a PHI-compliant GPT-5 deployment, backed by 40 hours of manual validation to control for annotator error.

The result the authors highlight is what they name an 'observability asymmetry.' Phase structure, the organized shape of a clinical visit, is recoverable from the transcript alone. Patient state, operationalized as PROM scores for voice, cough, and swallowing, is only partially observable, 'even in a setting designed to elicit patient symptoms and experiences.'

That distinction cuts against a common assumption underlying the current wave of clinical AI, from ambient scribes to encounter-quality scoring: that a well-annotated transcript is a reasonable proxy for the encounter itself. Structural things, whether the visit hit its expected phases, whether the clinician moved through a checklist, appear to travel through text. The lived symptom load a patient walked in with does not, at least not fully, even when the visit was designed to draw those symptoms out.

The scope of the claim is narrow and the authors flag it that way. It is one paper, one specialty, three symptom domains, and PROM scores as the yardstick for patient state. The abstract does not name which transcript features correlate with the parts of state that are observable, whether adding audio prosody or exam findings would close the gap, or whether the same asymmetry holds in primary care or mental health. What it does say plainly is a caution against 'transcript-only inference of human state.'

For health systems shopping AI scribes and encounter-analytics products, the practical read is: keep validated PROMs in the loop, and treat transcript-derived signals as good for workflow structure and coverage rather than as a stand-in for how the patient is actually doing.

Shared on Bluesky by 2 AI experts