Paper Uses Off-Policy Environment Evolution to Lift Qwen3.6-35B 18 Points on Terminal-Bench 2.1
Summary
Researchers introduce a training method that incrementally evolves terminal-agent environments off-policy generation by generation, validated against frontier models Hy4, Claude Opus 5 and GPT-5.6 Sol. Applied to smaller open models, the approach lifts Qwen3.6-27B by 14.4 points and Qwen3.6-35B-A3B by 18.0 points on Terminal-Bench 2.1. A multi-agent framework implements three evolution directions derived from multi-turn learning objectives, addressing prior co-evolution methods that stalled as models grew stronger.
Originally reported by arxiv.org
Read the original article →Original headline: Paper Uses Off-Policy Environment Evolution to Lift Qwen3.6-35B 18 Points on Terminal-Bench 2.1