huggingface.co web signal

Harness-Zero Paper Distills Agent-Harness Behaviors Into Weights, Lifts Base LLMs From 23.3% to 44.3%

Summary

A Sept 21 arXiv paper from Peking University introduces Harness-Zero, an 'agent-as-harness' training loop that distills behaviors from optimized coding/agent harnesses back into a student model's weights. On SpreadsheetBench Verified, AppWorld and USPTO Retrosynthesis, removing the harness after distillation lifts the base model from 23.3% to 44.3% macro-average - beating the 41.7% achieved when the specialized harness stays attached. The agent-as-harness variant also outperforms code-as-harness distillation 81.1% vs 78.1% across six benchmark-model settings, recovering 82.3% of harness-induced behaviors.