paper web signal

Occamy-1.0 open 35B edges GPT-5.6 Sol on Claw-Eval average

TL;DR

  • Occamy-1.0 averages 82.20 on Claw-Eval, just ahead of GPT-5.6 Sol at 81.80, built on the Qwen3.6-35B-A3B checkpoint.
  • On WildClawBench and AutomationBench Pass1, Occamy trails GPT-5.6 Sol (49.16 vs 67.20; 27.60 vs 45.50).
  • The paper releases model weights and a subset of training data, framing Occamy as "the low-cost knee" of a co-work cost-performance frontier.

On Claw-Eval, one of four representative benchmarks the paper reports on, Occamy-1.0 posts 82.20 on the average metric, a hair above the 81.80 it attributes to GPT-5.6 Sol. It comes from further training the Qwen3.6-35B-A3B checkpoint, and its weights along with a subset of the training data have been released.

The abstract makes a modest claim, not a crown one. "Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks," the authors write, placing the model at "the low-cost knee of the observed cost--performance Pareto frontier" under their own stated evaluation and pricing protocol.

The picture across the reported benchmarks is uneven. On WildClawBench, Occamy scores 49.16 against GPT-5.6 Sol's 67.20. On AutomationBench Pass1, 27.60 against 45.50. On Claw-Eval's pass3 metric it hits 71.40, ahead of GPT-5.6 Sol at 68.90 but behind Qwen3.8-Max at 73.70 and DeepSeek V4 Pro at 74.50.

The paper is candid about why cost matters. "Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered."

The abstract publishes no pricing assumptions behind the low-cost knee claim, and no GDPval score for Occamy appears in the head-to-head table pulled from the paper's HTML, though GDPval is one of the four representative benchmarks the paper cites.