arxiv.org web signal

Paper Uses Off-Policy Environment Evolution to Lift Qwen3.6-35B 18 Points on Terminal-Bench 2.1

Summary

Researchers introduce a training method that incrementally evolves terminal-agent environments off-policy generation by generation, validated against frontier models Hy4, Claude Opus 5 and GPT-5.6 Sol. Applied to smaller open models, the approach lifts Qwen3.6-27B by 14.4 points and Qwen3.6-35B-A3B by 18.0 points on Terminal-Bench 2.1. A multi-agent framework implements three evolution directions derived from multi-turn learning objectives, addressing prior co-evolution methods that stalled as models grew stronger.