611-Case Tumor Board Benchmark: Best of 14 Models at 3.43/5 Clinical Equivalence

Found first: a primary source the press has not covered yet.

A benchmark built from 12,534 minutes of real multidisciplinary tumor board recordings now gives a concrete measure of where AI stands against actual clinical specialists. The best of 14 tested models scored 3.43 out of 5 on clinical equivalence and 2.78 out of 5 on board consensus alignment. The dataset, leaderboard, and curation pipeline are described and publicly released in the paper on arXiv.

What the source says

The paper, by Anqi Li, Zhixuan Ge, Yixuan Duan, and nine co-authors, introduces OpenTumorBoard: 611 patient cases and 19,157 discussion turns drawn from 219 publicly recorded YouTube tumor board meetings spanning 12,534 minutes, covering 10 specialist roles. Two evaluation settings are defined. SPECIALIST TURN asks a model to answer a real clinician question mid-discussion. BOARD SIMULATION requires generating a full board discussion and reaching consensus on therapy recommendations, surgical plans, next actions, and clinical trial matching. Fourteen general-purpose and medical LLMs were tested across both settings. Three M.D. experts reviewed a subset and confirmed high information coverage and strong fidelity of the extracted consensus conclusions. Supervised finetuning and reinforcement learning both improved held-out test performance.

Why it matters

Most medical AI benchmarks use curated exam questions or synthetic cases, not the real-time consensus work of actual specialist teams. OpenTumorBoard derives its evaluation standard from what clinicians in session actually said and decided, which makes the scores harder to inflate through benchmark-specific training. The top clinical equivalence score of 3.43/5 quantifies a gap from specialist-level performance on questions those specialists actually posed to one another. On the harder BOARD SIMULATION task, the best model's board consensus alignment falls to 2.78/5. Supervised finetuning and reinforcement learning both improved held-out performance.