StarHarness lifts enterprise agents 20-35pp with no retraining
TL;DR
- StarHarness reports 20-35 percentage point gains over a default harness across three enterprise benchmarks, without model retraining.
- Just 4 to 12 accepted changes per environment produced the lift, and the improvements persist on tasks held out from the evolution search.
- The paper says the gains transfer across GPT and Qwen model families without re-optimization.
Harness edits alone produced 20 to 35 percentage point gains over a default harness across three enterprise agent benchmarks, according to a preprint from the StarHarness team, with no changes to model weights.
The authors, Esakkivel Esakkiraja and colleagues, describe a stratified search over prompts, task framing, tool interfaces, skills, and agent architecture. Some tasks are kept hidden from the search process and others reserved for generalization testing.
The paper reports the gains came from '4-12 accepted changes per environment' and persisted on tasks excluded from the evolution process. The improvements also transferred across GPT and Qwen model families without re-optimization.
The authors attribute the lift to interface corrections, environmental conventions, and compressed operational knowledge, which they say produced fewer false diagnoses and shorter execution trajectories.
Originally reported by paper
Read the original article →Original headline: StarHarness Adds 20–35pp on Enterprise Agent Benchmarks With Harness Changes Alone, No Retraining