Kaur: red-team benchmarks fall short on rare catastrophic harms
TL;DR
- Bandana Kaur's paper derives a closed-form 'evidential ceiling' bounding how much a single benchmark null result can shift belief about a model's safety.
- An audit of eight evaluation suites finds current benchmarks adequate for high-frequency harms but 'several orders of magnitude short' for rare catastrophic ones.
- The bound also covers adaptive and automated red-teaming, since discrimination between the hypotheses, not attack success, determines evidential worth.
For a couple of years the standard shorthand for AI safety has been to point at a benchmark score. A short new paper from Bandana Kaur, posted to arXiv in late July, gives that habit a formal audit and comes back with an awkward result.
Kaur defines an evidential ceiling for red-team evaluations, the largest factor by which one benchmark result can move belief under a fixed testing budget, and derives it in closed form for the benchmark null result. That produces a two-regime picture. Above a calculable harm rate, a benchmark of modest size can certify a category to a stated evidentiary standard, and a clean sheet becomes the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, 'no passive benchmark of feasible size provides the specified evidence of safety.' When she audits eight evaluation suites against the boundary, current benchmarks come out adequate for high-frequency harm categories and 'several orders of magnitude short for rare, catastrophic ones.'
The part that will land hardest with anyone building or regulating frontier models is that the bound is not specific to passive benchmarks. Written in terms of a procedure's hypothesis-conditioned elicitation rates, it also covers adaptive and automated red-teaming, and the paper's claim is that what generates evidential worth is discrimination between the hypotheses being tested, not attack success. That reframes a lot of the current investment in cleverer attackers.
The honest caveat is that this is a single-author preprint and the abstract does not name the eight audited suites or the specific evidentiary standard used, so treat the specifics as posed rather than independently confirmed, and expect scrutiny of the independence assumptions that make the closed form work.
If the result holds up, though, it hands regulators, labs and auditors a shared calculator for which safety claims a red-team result can actually support. Certification pressure on frontier models is only going up, and the labs and auditors who can state the harm-rate regime their tests cover, and the one they don't, will be in a much stronger position than the ones still waving aggregate benchmark scores at policymakers.
Originally reported by paper
Read the original article →Original headline: Closed-Form Proof: Red-Teaming Benchmarks Are 'Orders of Magnitude Short' for Rare Catastrophic Harms