paper web signal

Paper: better safeguard scores don't prove a safer LLM system

TL;DR

  • The paper argues safeguard evals reporting refusal, attack success and policy violation rates don't answer whether a deployed system is actually safer.
  • Evidence is asymmetric: one successful attack proves residual harm; proving safety needs system-level evidence, which the authors say is rarely produced.
  • Wu, Zhang and Yu conclude a better local score is not, by itself, a stronger claim about the deployed system.

Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. A paper posted to arXiv on September 1 by Pingyu Wu, Weiming Zhang and Nenghai Yu argues those numbers do not answer the question a deployment actually has to answer.

"A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in," the abstract states.

The central move is what the authors call an asymmetric evidence requirement. "One attack that obtains harmful help from the deployed service suffices to establish that such help remains," the abstract says, and the authors note that such attacks "appear repeatedly in the coded record." Establishing the opposite, that little residual harm remains, cannot come from the safeguard's own numbers. It also requires evidence about what the surrounding system still allows after the safeguard performs its local function.

That kind of evidence, they write, is "supported or derived in only a small minority of the depth-coded claims," and only one such claim bounds its scoped residual. From this the authors draw a blunt conclusion: "A better local score is therefore not, by itself, a stronger claim about the deployment."

The abstract does not name a framework or publish a per-benchmark scoreboard. It stops at the argument. "Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer."