arxiv.org web signal

Feng paper: LLM cross-state Medicaid errors look overstated

TL;DR

  • A pre-registered test fixes the question wording and varies only the jurisdiction across 50 states plus DC, for three Medicaid income-eligibility quantities.
  • Claude Sonnet 5.5 reproduced another state's current value in 10 of 153 items; GPT-5.6 Sol did so in 25 of 153.
  • Crediting any wrong answer that matches another state yields 3-5x more apparent substitutions than checking the asked state's own full records.

A new arxiv preprint by Jiayu Feng takes aim at a specific failure mode: when a language model answers a state-specific policy question wrongly, is it hallucinating, or is it returning a real value that happens to hold in a different state?

The design is deliberately narrow. The question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia, for three exactly defined Medicaid income-eligibility quantities, 153 items in total. Gold values come from an official data book that agreed with an independent source on 101 of 102 checked cells. Under a pre-registered protocol with two independent repeats, Claude Sonnet 5.5 reproducibly returned another state's current value for 10 of 153 items; GPT-5.6 Sol did so for 25. Two researchers we track posted the preprint the same day.

The real finding is about attribution. "Crediting any wrong answer that equals another state's value yields 3-5x more reproducible substitutions than checking every number in the asked state's own records," the paper reports, because many apparent cross-state answers are "the asked state's own values under another convention or from an earlier year." Feng's conclusion is flat: "Claims about cross-jurisdiction error need a complete same-state reference set."

Shared on Bluesky by 2 AI experts