Schulman Panel Casts Doubt on Imminent Recursive Self-Improvement
TL;DR
- Horizon length on agent tasks is doubling every three months, per Baseten's Charlie O'Neill, though he attributes it to horizon generalization, not cross-domain reasoning.
- John Schulman of Thinking Machines says defining objectives will remain the 'last job for humans' as AI capabilities scale.
- Zyphra CTO Beren Millidge argues most gains attributed to reinforcement learning actually come from 'very, very good mid-training data'.
The rate at which AI models can work on a task before falling over is "doubling every three months," Charlie O'Neill, Head of Model Training at Baseten, told Dwarkesh Patel in a roundtable published September 11. What that curve does not do, he argued, is prove models are learning to reason across domains. It shows "horizon generalization…how to continue on that task for longer."
O'Neill was on the panel with John Schulman, chief scientist at Thinking Machines and a former OpenAI co-founder who led its RLHF work, and Beren Millidge, CTO of the open-source lab Zyphra. The three spent the episode weighing whether recursive self-improvement is a near-term event or a longer story.
Schulman's caution was behavioral first. "A new model comes out and people are blown away…but then they use it a bit, and it starts to feel dumb after a month," he said. On the widely floated idea that frontier models will learn from live enterprise deployment, he flagged a plain constraint: "Companies aren't going to want to have the model provider learn from all of their deployment."
Millidge pushed back on the notion that reinforcement learning is doing the heavy lifting in current gains. "An awful lot of what we see as successes of RL actually comes from very, very good mid-training data," he said. He allowed that architecture still matters at the margin: architectures, he said, "let you reach a qualitatively new regime which you couldn't reach."
O'Neill's structural worry was starker. Continuing improvement, he suggested, may require a discovery current methods cannot make: "If it requires another one of those discontinuities to solve, I'm not sure that the current method of training LLMs…would be able to discover that discontinuity."
Schulman closed on where humans still fit. "The last job for humans…is defining the objective and deciding what we actually want," he said. That framing lands in a week when the safety debate has been running hot in our tracker; the episode does not settle it, but three practitioners this close to the training loop treating objective specification as the binding constraint is itself the news.
Originally reported by dwarkesh.com
Read the original article →Original headline: Dwarkesh Panel With Schulman, Millidge and O'Neill Debates Whether Recursive Self-Improvement Is Imminent