paper web signal

WearableQA benchmark stretches 14 LLMs from 19.6% to 72.9%

TL;DR

  • WearableQA comprises 4,084 ten-option multiple-choice questions built from 200 real users' wearable, blood biomarker, and demographic data, with up to 500 days each.
  • Across 14 proprietary and open-source LLMs, scores span 19.6% to 72.9% against a 10% chance baseline.
  • Most models score below 60%, and the authors conclude the benchmark remains far from solved.

Scores across fourteen large language models spread from 19.6% to 72.9% on WearableQA, a benchmark posted to arXiv on Sept. 4 that tests health reasoning against real users' longitudinal wearable records rather than synthetic or clinic-curated data.

The set comprises "4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements." Sixteen question types split along two axes: data versus health reasoning, and single- versus cross-signal reasoning.

To keep the questions honest, Ji Soo Lee and seven co-authors used a "dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns." The paper reports that "most models achieve accuracies below 60%," and the abstract does not disclose which system hit 72.9% or which sat at the 19.6% floor.