paper web signal

OpenTumorBoard: top LLMs hit 3.43/5 on real tumor boards

TL;DR

  • The benchmark covers 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of YouTube tumor board recordings.
  • The best of 14 frontier and medical LLMs tested scored 3.43 out of 5 on clinical equivalence and 2.78 out of 5 on alignment with recorded board conclusions.
  • Supervised fine-tuning and reinforcement learning on the training split yielded substantial improvements on a held-out test set, per the authors.

The best of 14 frontier and medical LLMs tested on real multidisciplinary tumor boards scored 3.43 out of 5 for clinical equivalence to specialist answers, and 2.78 out of 5 for alignment with the board's recorded conclusions. The OpenTumorBoard paper on arXiv builds the benchmark from 611 patient cases and 19,157 turns of cancer discussions across ten specialist roles, transcribed from 12,534 minutes of publicly recorded video on YouTube.

The dataset spans 219 recorded meetings and 16,215 questions put to specialists during the sessions. Three M.D. experts reviewed the cases, confirming what the authors describe as "high information coverage and factuality of patient cases."

The authors position OpenTumorBoard as a training resource as well as a measuring stick, reporting that "supervised finetuning and reinforcement learning yield substantial improvements on a held-out test set." They frame the headline scores as revealing "substantial limitations" in current models' capacity for multidisciplinary cancer decision-making.