arxiv.org web signal

Paper: API Benchmark Scores Overstate Chatbot Accuracy by 3.4pp

TL;DR

  • Auditing ChatGPT, Claude and Gemini across seven systems and nine benchmarks, API evaluations averaged 3.4 percentage points higher in accuracy than chatbot interfaces.
  • For ChatGPT, the API-versus-interface performance gap exceeded the API-only difference between GPT 5.3 and GPT 5.4, roughly a full model generation.
  • Adjusting system prompts, sampling parameters and reasoning settings shifted behavior in some cases but did not reliably close the gap, the authors report.

API benchmark scores overstate what users get inside the deployed chatbot, according to a new arXiv preprint by Jennifer Wang, Joachim Baumann, Daniel E. Ho and Sanmi Koyejo. Auditing ChatGPT, Claude and Gemini across seven systems and nine benchmarks spanning general capability, social bias and sycophancy, the authors report API evaluations "score 3.4 percentage points higher in accuracy" and 2.1 percentage points higher in test-retest consistency than the same models accessed through their chatbot interfaces. Two experts in our Who's Who directory shared the paper.

The ChatGPT gap crosses generation lines. "The performance difference between API and interface access exceeds the API-only difference between GPT 5.3 and GPT 5.4," the paper reports; put another way, "switching access surfaces can degrade performance as much as downgrading a full model generation."

Whether the usual API controls can bridge the divide gets its own test. Varying system prompts, sampling parameters and reasoning settings "shift behavior in some cases but do not reliably eliminate the gap." The authors label the underlying problem a "context-validity gap," where "measurements obtained through APIs do not necessarily generalize to corresponding deployed interfaces."

The abstract publishes averages but no per-benchmark or per-system numbers, and does not say which of the seven pipelines produced the widest divergence.

Shared on Bluesky by 2 AI experts