huggingface.co web signal

Hugging Face paper ships 50,228 agent error-diagnosis pairs

TL;DR

  • The Agent Error Dataset (AED) collects 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models.
  • On 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a 32.7 percentage-point gain.
  • Fine-tuning Qwen3-8B on the diagnosis task lifts exact-step agreement from 47.2% to 63.6% on a 943-case holdout, past a 54.7% prompted baseline.

50,228 error-diagnosis pairs from 9,961 agent tasks now sit in a public corpus released alongside a new Hugging Face paper, built for post-training LLM agents on their own failures rather than only on successful trajectories. The dataset, called Agent Error Dataset (AED), spans 33 environments, 19 harness families and 23 policy models.

The authors (Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li and Heng Ji) open on the premise that "An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses." A five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence.

Across 3,062 matched replay pairs, the paper reports that "first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points." On a 943-case holdout, fine-tuning Qwen3-8B on the diagnosis task takes exact-step agreement from 47.2% to 63.6%, past the 54.7% a prompted reference baseline reaches. On WebShop-lite, training the model to propose action-only repairs scored 6.67 percentage points above success-only training.

The abstract does not state AED's license, name the specific 19 harness families or 23 policy models included, or report how the gains transfer to agent stacks the authors did not sample. The release lands in a thick 417-story run of agent coverage we have logged in the last 90 days.