⚡ 26 h early
automated discovery no universal harness paper
Automated Discovery Has No Universally Superior Harness
2 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
“Paper : arxiv.org/abs/2607.18235 Github : github.com/akshat57/har...”
“Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen Automated Discovery Has No Universally Superior Harness https://arxiv.org/abs/2607.18235”
2 experts discussed this · 15 posts
Leshem (Legend) Choshen @EMNLP: OpenEvolve underperforms simple autmated discovery harnesses the rest are insignificant from each other. The best choice changed across model–problem pairs. We ran a controlled study (3m+ rollouts)…
Leshem (Legend) Choshen @EMNLP: Automated discovery has high run-to-run variance, yet harnesses are often evaluated with only 3–5 runs. When a harness performs better, how do we know it is genuinely better—and not simply lucky?
Leshem (Legend) Choshen @EMNLP: To answer this, we systematically evaluated 30 budget-matched harnesses across 12 model–problem pairs, using repeated-trial statistical analysis. We found :
Open the full discussion →