Riedl, Matlin: no LM wargame fit for policy without audit
TL;DR
- Mark Riedl and Glenn Matlin argue no LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case.
- The paper names five failure modes: decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination.
- Ordinary benchmarks cannot establish safety for these settings; wargames can expose failures as stress tests but are not themselves safety cases for consequential use.
No language-model-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case. That is the central claim of a position paper by Mark Riedl and Glenn Matlin, submitted to arxiv on 31 July 2026 and presented at the Social Sim'26 Workshop at COLM 2026.
The authors name five failure modes: decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination. The core worry, in their words: "model language determines both what an actor attempts and what becomes simulated reality." The same features that make LMs attractive here — playing agents, generating scenario branches, adjudicating ambiguous actions, summarizing lessons — are exactly what make them dangerous when the output is read as evidence.
"Ordinary benchmarks cannot establish safety for these settings," the paper argues. The proper use of open-ended wargames today, Riedl and Matlin contend, is to stress-test decision-influencing LM agents, not to stand in as a safety case themselves. The abstract does not name a specific vendor, model, or government program the warning is aimed at.
Shared on Bluesky by 1 AI expert
Originally reported by arxiv.org
Read the original article →Original headline: Position: AI Is Not Ready for Strategic Conflicts