arxiv.org web signal

Wallach paper reframes GenAI evaluation as measurement science

TL;DR

  • A NeurIPS 2024 workshop paper led by Hanna Wallach argues evaluating generative AI systems mirrors measurement problems long studied in the social sciences.
  • The framework separates four levels: background concept, systematized concept, measurement instrument, and instance-level measurements.
  • The authors argue ML researchers typically jump from background concept straight to instrument, with little explicit systematization in between.

There is a quiet argument going on in AI evaluation right now about whether the field even knows what it is measuring, and a NeurIPS 2024 workshop paper led by Hanna Wallach with nineteen collaborators is one of the sharper statements of it. The claim is that measuring things like a generative model's capabilities, impacts, opportunities, and risks is not really a novel technical challenge. It is the same measurement problem social scientists have been chewing on for decades, and machine learning has been skipping a step.

The framework the authors propose has four levels: the background concept, the systematized concept, the measurement instrument or instruments, and the instance-level measurements themselves. Their diagnosis of the current state of ML evaluation is blunt. Researchers and practitioners, they write, "appear to jump straight from background concepts to measurement instruments, with little to no explicit systematization in between." In plain terms, someone decides they want to measure "reasoning" or "harm" or "helpfulness," and the next thing that happens is a benchmark exists. The step where you pin down exactly which version of the fuzzy concept you are actually operationalising, and why, tends to get skipped.

Why that matters for anyone building on top of these models is the validity question. If the systematization step is invisible, so are the assumptions baked into the score. A model that tops a benchmark for a background concept called "safety" is only as trustworthy as the systematized version of "safety" that the benchmark actually captures, and right now that middle layer is usually implicit. The authors pitch their framework as a set of lenses for interrogating that validity, and as a way to let stakeholders from outside ML take part in the conceptual debate rather than only the leaderboard one.

The honest caveat is that this is a position paper from a workshop track, not a new benchmark or a tool. It does not tell you which of today's popular evals fail its validity lenses, and it does not prescribe how a fast-moving product team should retrofit the four levels into an existing release pipeline. Whether major labs actually adopt the systematization step, or whether the vocabulary gets absorbed without the underlying work, is the part that will decide if this lands.

The forward-looking read is that policy teams, auditors, and social scientists finally get a shared grammar with ML researchers for asking "what does this number actually mean," and that is the conversation the next round of AI regulation is going to need.

Shared on Bluesky by 1 AI expert