ActiveSaddler Adds 7.5pp on Terminal-Bench via Bandit Curriculum
TL;DR
- ActiveSaddler improves test Pass@1 by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 over a fixed-curriculum baseline.
- The method models agent training as a non-stationary bandit whose arms are reusable failure patterns that update as the harness evolves.
- Ablations tie the gains to doing all three at once: dynamic targets, utility estimation, and exploration of unseen failures.
ActiveSaddler, a method for training large-language-model agents, raises test Pass@1 by 4.4 percentage points on GAIA2 and 7.5 points on Terminal-Bench 2.0 over the same harness optimizer run with a training-scenario order fixed before optimization, the authors report in a paper posted to arXiv.
The twist is to let the curriculum move. Existing harness-optimization methods, the authors write, "primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates." ActiveSaddler instead "models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets," grouping recurring agent failures into reusable failure-pattern arms and reweighting them as the harness improves.
Ablations attribute the gains to three ingredients working jointly: "dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery." The authors frame the result as establishing "automated curriculum learning as a new crucial optimization dimension for harness optimization."
The abstract reports no absolute Pass@1 numbers for either benchmark, no comparison against other curriculum-learning baselines, and no indication of which base model the agent harness runs on top of.
Originally reported by paper
Read the original article →Original headline: ActiveSaddler Adds +7.5pp on Terminal-Bench by Treating Agent Training Curriculum as a Bandit