arxiv.org web signal

Google, DeepMind team pushes for longitudinal AI evaluations

TL;DR

  • Google Research and DeepMind authors propose replacing single-prompt AI evaluations with long-term measurements of user behavior over sustained interactions.
  • The paper flags loneliness, cognitive deskilling, ideological polarization and emotional dependence as risks that surface only after extended use.
  • It sketches three data pipelines, RCTs, field studies and user simulations, but stops short of proposing a working measurement stack.

The interesting AI safety papers this year keep pointing at the same gap: we evaluate models on their answers to a single prompt, but people use them for months. A new preprint from Google Research and Google DeepMind on arXiv argues that whole framing needs to change, because the risks worth caring about, like loneliness, cognitive deskilling, ideological polarization, and emotional dependence, only show up after sustained use.

The authors, Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton and Roma Patel, make the case that language models occupy a strange category because of their 'human-ness' and rapid integration into users' daily lives. That combination, they write, can introduce 'longitudinal risks,' which they describe as 'cognitive, developmental and socio-affective changes in humans' that might not surface during a short-term interaction. Their proposed fix is a pivot: away from static benchmarks of text generation and toward long-term measurements of behavioral change in the humans on the other side.

Practically, the paper sketches three data-collection routes. Randomized controlled trials give you causal isolation but suffer attrition. Field studies give you ecological validity through naturalistic use. Simulated users let you look forward at risks but drift over long trajectories. On top of that it borrows heavily from social science, arguing that 'psychometric scales are designed to measure psychological constructs at the level of individual users,' and pairs them with computational methods drawn from Rational Speech Acts frameworks and dynamic systems theory to identify behavioral inflection points in long conversations.

What the paper does not give you is a working measurement stack. It reads as a position paper, not an empirical result, and the risk categories it names are ones the field has been discussing publicly for a while. The specific validated instruments, and the thresholds that would trigger an intervention, are still hand-waved. Take it as an agenda from a large lab, not a benchmark you can run.

The interesting part for anyone building consumer AI is what happens if this framing lands. The next round of alignment work would be judged on what happens to users over weeks and months, not on a single-turn eval leaderboard. That is a very different measurement problem, and the labs that get their telemetry and their psychometric partnerships ready for it first will have a real head start.

Shared on Bluesky by 2 AI experts