paper web signal

CompoWorld's 35B agent beats Claude Opus 4.6 on AutomationBench

TL;DR

  • CompoWorld builds a sandbox of 448 services and 10,130 tools that chain across services to generate verifiable multi-step agent tasks.
  • The authors post-train Qwen3.6-35B-A3B on 3,000 supervised fine-tuning trajectories and 1,000 reinforcement-learning tasks distilled from the environment.
  • The resulting model reportedly surpasses Claude Opus 4.6, with a 9.17-point average gain across eight benchmarks.

A 35-billion-parameter open model, trained on tasks generated inside a sandbox of 448 services and 10,130 tools, reportedly beats Claude Opus 4.6 on AutomationBench. That is the central claim of CompoWorld, a preprint by Xiao-Wen Yang and colleagues.

The framework "expands the task space by composing a finite library of reusable services," the abstract says, chaining tools across services to produce verifiable multi-step tasks. From that library the authors generate 3,000 supervised fine-tuning trajectories and 1,000 reinforcement-learning tasks, then post-train Qwen3.6-35B-A3B on the mix.

The trained model "surpasses frontier models such as Claude Opus 4.6," the paper reports, with a "9.17 points average improvement across eight benchmarks."

The abstract does not name those eight benchmarks, and it publishes no per-benchmark deltas or specific Opus 4.6 scores. Whether the 448 services wrap real APIs or are self-contained mocks is also not stated. That distinction would matter for anyone trying to reproduce the recipe on their own tool stack.