MicroVerse probes identity drift in multi-agent LLM simulations
TL;DR
- In a 50-by-50 grid with scarce water, agents most often added self-deception resistance as a new boundary: 27 of 111, or 24%.
- The framework pairs an immutable core identity file with a mutable working identity and a three-layer memory architecture that triggers periodic reflection.
- Authors call the results preliminary existence proofs, run on a single model with n=25 agents per condition.
When 25 language-model agents were dropped into a 50-by-50 grid with scarce water and asked to keep track of who they were, the boundary they most often invented for themselves was resistance to self-deception, accounting for 27 of 111 added boundaries, or 24%, according to the MicroVerse preprint on arXiv.
The instrument, from Sky Ng, Brihi Joshi, Ishan Gupta, Shirley Huang and 41 co-authors, hands each agent an immutable core identity file alongside a mutable working identity, and uses a three-layer memory architecture so agents periodically compare who they now are against who they started as. An eight-action verb space maps every move to a moral boundary, and uniform longitudinal snapshots are taken every N ticks to sidestep survivor bias. Two of the AI researchers we follow circulated the link the day it landed.
The self-deception result surfaced unprompted, a category the authors did not seed. Robustness checks across reflection-threshold conditions of 40, 80 and 150 showed that lower thresholds "accelerate and increase revision frequency but preserve drift direction," per the abstract.
The authors describe the runs as "preliminary existence proofs" rather than statistically significant claims: a single model, n=25 agents per condition, no cross-family comparisons in the retrieved abstract.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations