Builder LLM Lifts Frozen Target 8.95 Points via Meta-Skills
TL;DR
- A Builder LLM learned reusable 'meta-skills' that lifted a frozen Target model by 8.95 points on Harness-Bench and NewtonBench macro-averages.
- Handing the same skill bank directly to the Target, without the Builder's scaffolding step, performed 12.02 points worse than Builder-mediated application.
- Gains persisted when one model played both Builder and Target, which the authors frame as a route to system-level self-improvement without weight updates.
Freeze both the model doing the task and the model setting up its workspace, teach the second one reusable principles for when and how to help, and the first gets 8.95 percentage points better. That is the headline result in a new preprint on arxiv from Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang and Heng Ji, submitted September 29.
The authors call these principles 'Meta-Skill' and define them as 'principles specifying when support is needed and what resources to provide.' A Builder model observes a Target's execution feedback on a development set, distills what worked into a frozen skill bank, then uses that bank to construct harnesses for unseen tasks.
Across Harness-Bench and NewtonBench, full-bank meta-skills beat no-skill construction by 8.95 macro-average points, and beat direct delivery of the same bank to the Target by 12.02 points. The abstract notes gains persist 'when the same model serves both roles,' which the authors read as 'a path to system level self-improvement through learning to build better environments.'
The abstract publishes no per-benchmark breakdown, no base model identity, and no absolute scores.
Originally reported by paper
Read the original article →Original headline: AI Builder Learns Reusable Meta-Skills to Scaffold a Separate AI—+9 Points, No Weight Updates