technologyreview.com web signal

Ex-AlphaGo scientist: LLMs don't reason, just predict tokens

TL;DR

  • Thore Graepel, former core member of the AlphaGo team at DeepMind, argues contemporary LLMs lack an explicit, persistent, and inspectable epistemic state.
  • He contrasts AlphaGo's split of neural hunches with explicit search deliberation, while LLMs fuse knowledge and reasoning inside the network's weights.
  • Chain-of-thought is extended next-token prediction, not System 2 deliberation, and models often concoct rationales after reaching an answer by another route.

Thore Graepel, chair of machine learning at University College London and a former core member of the AlphaGo team at DeepMind, argues in MIT Technology Review that today's large language models only look like they reason. The essay places chain-of-thought prompting firmly on the intuition side of the ledger, not the deliberative side.

Graepel reaches for Daniel Kahneman's framework. "A well-known theory in the behavioral sciences, popularized by Daniel Kahneman, distinguishes between two modes of human thought: System 1 is fast, gut-level, effortless; system 2, slow, step-by-step, and deliberative," he writes. AlphaGo, he says, had both. "Its networks supplied the hunches—this move looks promising, this position looks won—and its search supplied the deliberation, testing those hunches against the moves and countermoves that would follow." A language model, by contrast, "picks the next token, over and over."

That matters because chain-of-thought is often sold as deliberation. Graepel says it is not. "The intermediate reasoning is still produced by the same next-token prediction process, iterated for longer before the model commits to an answer." He lists three shortcomings. Models maintain "no explicit, persistent, and inspectable epistemic state." Knowledge and reasoning are entangled in the weights. And when models do produce rationales, "research has demonstrated that the bots often concoct them after the fact, reaching an answer by one route but reporting another."

The comparison to classical game AI is pointed. Deep Blue beat Garry Kasparov in 1997 by looking "six to eight moves ahead per player and evaluating 200 million chess positions per second, using rules hard-coded by humans." AlphaGo's Move 37 in game two of its five-game match against Lee "looked so absurd that some commentators thought it was a programming glitch," then proved prescient. "I thought AlphaGo was based on probability calculation and that it was merely a machine," Lee said afterwards.

Graepel's prescription is architectural, not more scaling. "I believe we need a fresh approach to machine reasoning—one that draws on AlphaGo's architecture," he writes, with "reasoning... understood as a sequence of moves that change the epistemic state to advance knowledge and reduce uncertainty."

Shared on Bluesky by 1 AI expert