arxiv.org web signal

HELM Benchmarks 30 Language Models on 42 Scenarios, 7 Metrics

TL;DR

  • HELM benchmarks 30 open, limited-access, and closed language models across 42 scenarios and seven metrics under standardized conditions.
  • The seven metrics are accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency — measured for 87.5% of core scenario–model pairs.
  • Average coverage of the core scenarios rose from 17.9% before HELM to 96.0%, and 21 of 42 scenarios were new to mainstream LM evaluation.

Before HELM, prominent language models were compared on an average of just 17.9% of a common set of evaluation scenarios; the paper notes that some prominent models did not share a single scenario in common. Stanford's Holistic Evaluation of Language Models benchmarks 30 open, limited-access, and closed models across 42 scenarios and seven metrics — accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency — and raises that coverage figure to 96.0%.

The multi-metric design is the deliberate move. Measuring seven metrics for each of 16 core scenarios, the authors write, "ensures metrics beyond accuracy don't fall to the wayside, and that trade-offs are clearly exposed." Twenty-one of the 42 scenarios "were not previously used in mainstream LM evaluation," pushing coverage into areas the paper flags as underrepresented, including "question answering for neglected English dialects" and "metrics for trustworthiness."

The evaluation surfaces 25 top-level findings and releases all raw model prompts and completions publicly, alongside a modular toolkit the authors describe as "a living benchmark for the community, continuously updated with new scenarios, metrics, and models."

Shared on Bluesky by 1 AI expert