CompoWorld's 35B agent beats Claude Opus 4.6 on AutomationBench
TL;DR
- CompoWorld builds a sandbox of 448 services and 10,130 tools that chain across services to generate verifiable multi-step agent tasks.
- The authors post-train Qwen3.6-35B-A3B on 3,000 supervised fine-tuning trajectories and 1,000 reinforcement-learning tasks distilled from the environment.
- The resulting model reportedly surpasses Claude Opus 4.6, with a 9.17-point average gain across eight benchmarks.
A 35-billion-parameter open model, trained on tasks generated inside a sandbox of 448 services and 10,130 tools, reportedly beats Claude Opus 4.6 on AutomationBench. That is the central claim of CompoWorld, a preprint by Xiao-Wen Yang and colleagues.
The framework "expands the task space by composing a finite library of reusable services," the abstract says, chaining tools across services to produce verifiable multi-step tasks. From that library the authors generate 3,000 supervised fine-tuning trajectories and 1,000 reinforcement-learning tasks, then post-train Qwen3.6-35B-A3B on the mix.
The trained model "surpasses frontier models such as Claude Opus 4.6," the paper reports, with a "9.17 points average improvement across eight benchmarks."
The abstract does not name those eight benchmarks, and it publishes no per-benchmark deltas or specific Opus 4.6 scores. Whether the 448 services wrap real APIs or are self-contained mocks is also not stated. That distinction would matter for anyone trying to reproduce the recipe on their own tool stack.
Originally reported by paper
Read the original article →Original headline: CompoWorld's 448-Service, 10K-Tool Sandbox Beats Claude Opus 4.6 on AutomationBench