Pezzulo team argues embodied AI needs grounded world models
TL;DR
- A group of computational neuroscientists led by Giovanni Pezzulo argues current AI leans on passive linguistic training rather than grounded environmental interaction.
- The paper points to five neural circuit families including navigation, affordance perception, active exploration, allostatic control, and self-versus-external distinction.
- The authors say embodied AI is missing intrinsic dynamics, action-centered learning, autonomous open-ended learning, and social grounding aligned with human norms.
A short position paper landed on arxiv this month from a group of computational neuroscientists that reads as a direct challenge to how most frontier labs are building toward embodied AI. Giovanni Pezzulo and co-authors argue, in a paper posted to arxiv on July 15, that current systems lean on passive training regimes where linguistic regularities create the scaffold, and that biological organisms do something very different, building knowledge through environmental interaction first, with language layered on top of that grounded foundation.
The interesting part is that they get concrete about what 'grounded' actually means. The paper points to five families of neural circuits that, in their view, are already doing the job in biological brains: navigation in physical and conceptual spaces, affordance-based object perception and interaction, active exploration and perceptual learning, allostatic control and emotion regulation, and the machinery that lets an organism distinguish self-generated from external outcomes. Each of those is a specific claim about how animals learn, not a metaphor for a training objective.
Against that list, the authors mark what they see as missing from today's embodied AI work: intrinsic dynamics as a foundation for learning, the centrality of action in aligning internal dynamics with external reality, autonomous open-ended learning prioritized over passive data assimilation, and early predictive mechanisms that scaffold higher cognition like reasoning and planning. Their proposed direction then adds social interaction on top, so that world models are not only grounded but also socially shared and aligned with human norms and values.
The honest caveat is that this is a position paper, not a benchmark result. It does not propose a specific architecture, does not put numbers on how far current systems fall short, and does not describe what it would cost to build a system along the lines it advocates. What the reporting doesn't give you is any empirical test of whether these circuits, translated into engineering, would actually close the gap.
For teams betting that scaling text and video is the road to embodied competence, though, this is the kind of counter-argument worth reading before the next training run, because the authors are pointing at ingredients like action, intrinsic dynamics, and social grounding that no amount of extra tokens will supply.
Shared on Bluesky by 3 AI experts
-
arxiv.org/abs/2607.13560
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: Grounded world models in biological organisms and future embodied AI