paper web signal

DCP audit logs zero recoveries across 96 AI-agent episodes

TL;DR

  • The Discovery Certification Protocol logged zero recoveries in 96 episodes across two controlled audits, in SQLite optimization and virtual catalyst control.
  • Gate 2 hands matched agents registered starting information and observed Web content while withholding the target research history.
  • Paired studies produced 30 truthful recoveries versus zero neutral recoveries, and a deterministic, LLM-free verifier reproduces the decisions from frozen evidence.

Zero recoveries in 96 episodes. That is the headline number from a September 7 arXiv preprint by Jingjie Ning, Shanshan Zhong, Xiaochuan Li and Ji Zeng, proposing what they call the Discovery Certification Protocol, or DCP, for auditing claims made by AI research agents.

The protocol turns AI-agent claims into 'executable recovery and feedback tests' spread across three gates. Gate 1 checks improvement on sealed evaluation. Gate 2 hands matched agents 'the registered starting information and observed Web content while withholding the target research history' and requires that no valid method reach the numerical target. An optional Gate 3 measures the average effect of truthful feedback against a specified neutral policy.

Two controlled audits carry the protocol through. One covered SQLite optimization; the other, virtual catalyst control. 'Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468,' the abstract reports. In paired follow-ups, 'each paired study yielded 30 truthful recoveries and zero neutral recoveries,' with passing 60-pair null studies.

The authors argue DCP amounts to 'a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research,' anchored by a deterministic, LLM-free verifier that reproduces decisions from frozen evidence. Both audit domains are narrow, and the abstract runs the results 'under different models' without naming any.