huggingface.co web signal

Opera 'Persistent Note' Critic Lifts Coding Agents up to 15pp

Agents Coding Tools Safety ai-research

TL;DR

  • Opera manages each critic diagnosis as a persistent note that stays open against a fixed resolution criterion, with admission and release audits gating it on both ends.
  • Across Terminal-Bench 2.1, SWE-Bench Pro and DeepSWE v1.1, Opera raises resolve rates by up to 12.4, 15.0 and 8.9 pp over four policy models.
  • Fine-tuning Qwen3.5-9B on Opera-guided student rollouts matches stronger-teacher distillation on held-out SWE-Bench Pro (+10.2 pp) while preserving Terminal-Bench 2.1 performance.

A critic framework called Opera lifts coding-agent task resolve rates by up to 15.0 percentage points across three repository-level benchmarks by treating every mid-run intervention as a persistent note that opens on a diagnosed issue and only closes when evidence satisfies a fixed resolution criterion. The paper on Hugging Face, posted October 9, frames verbal criticism as "an intervention whose value is revealed only afterwards" and builds around three questions: when the critic should intervene, whether the diagnosis is justified, and whether the fix actually took.

Opera combines a periodic review every k turns with event-driven triggers for Idle, Repeat, Error, Claim and Submission. Each delivered note names one of nine typed operators spanning the search, view, edit, test and submit stages, cites specific evidence, and ships with a resolution criterion that is checked by a separate release audit before the note is closed. An admission audit sits on the other end, rejecting unsupported or redundant proposals.

Across Terminal-Bench 2.1 (89 tasks), a 100-task SWE-Bench Pro subset and DeepSWE v1.1 (113 tasks), Opera posts the highest mean resolve rate against four baselines, SWE-PRM, SWE-Search, LLM-as-verifier and Agentic Rubrics, with Qwen3.8-27B as policy and GPT-5.6-Sol as default critic. The paper reports gains "by up to 12.4, 15.0, and 8.9 pp" on the three benchmarks across four policy models, with weaker Qwen3.5-9B drawing the largest lifts and already-strong agents gaining less. It also flags a recovery-disruption trade-off: stronger models convert feedback into more rescues but also regress more often on tasks they would otherwise solve.

The audits are load-bearing. Strip them out and the Terminal-Bench 2.1 gain collapses from 7.9 to 2.6 pp, with DeepSWE falling from 8.9 to 4.5 pp. Opera is one of 128 coding-tools stories we've tracked in the last 90 days, and the training-side result is where it earns most attention: fine-tuning Qwen3.5-9B on Opera-guided student rollouts matches distillation from a stronger model on held-out SWE-Bench Pro repositories (+10.2 pp) while preserving Terminal-Bench 2.1 at 25.8%, where distilling the Qwen3.8-27B teacher directly instead drops it from 23.6% to 7.9%. The authors attribute the difference to the student receiving "expert corrections at the states it actually visits."