datasets-benchmarks-proceedings.neurips.cc web signal

Raji and Bender challenge the 'general' AI benchmark myth

TL;DR

  • A NeurIPS 2021 position paper argues AI's most influential benchmarks are wrongly treated as broad, general measures of progress toward flexible systems.
  • The authors frame the problem as construct validity: state-of-the-art scores get read as milestones toward long-term goals they cannot actually measure.
  • The paper is by Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton and Alex Hanna, spanning Mozilla, UW Linguistics and Google Research.

A NeurIPS 2021 position paper that still gets cited whenever the industry debates what a leaderboard win actually means is worth revisiting on its own terms. AI and the Everything in the Whole Wide World Benchmark, by Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton and Alex Hanna, argues that AI's habit of treating a small set of influential benchmarks as stand-ins for progress toward general capability is a category error.

The abstract puts the case plainly. There is, the authors write, 'a tendency across different subfields in AI to valorize a small collection of influential benchmarks', and those benchmarks operate as stand-ins for 'a range of anointed common problems that are frequently framed as foundational milestones on the path towards flexible and generalizable AI systems'. The authors' target is not any single benchmark; it is the framing that reads state-of-the-art scores as evidence of progress toward long-term general capability. They call this a construct validity problem, borrowing a term from measurement theory: the benchmark may be well-defined, but what it measures is not the thing the community claims it measures.

That framing has become more consequential since 2021, not less. Model cards routinely lead with a bar chart of benchmark wins, and procurement decisions, regulatory conversations and investor pitches lean on those numbers as if they were IQ scores for systems. The Raji et al. paper is the position piece those bar charts eventually run into.

The abstract does not name the benchmarks the authors have in mind, does not propose a replacement, and stops short of the empirical case that the full paper makes; a reader who wants to know which specific evaluations the authors would retire has to open the full PDF on arXiv. It is a position paper, not a measurement study, and treating it as an off-the-shelf audit method would overstate what the abstract itself claims.

Still, the paper's usefulness is as a citation with teeth. Anyone commissioning an AI system for a specific, high-stakes context now has a peer-reviewed reason to insist that leaderboard wins are not the same as fitness for their task, and that the burden of proof for 'general' capability sits with the vendor making the claim.

Shared on Bluesky by 1 AI expert