aclanthology.org web signal

ACL paper: LLM political-bias evals shift when unforced

TL;DR

  • An ACL 2024 Outstanding Paper argues that LLM political-bias tests using multiple-choice surveys do not reflect how real users query models.
  • In a Political Compass Test case study, models gave substantively different answers when not forced into the test's fixed-choice format.
  • The authors also report that LLM answers shift depending on how models are constrained, and lack paraphrase robustness.

Multiple-choice surveys are a poor proxy for how people actually talk to LLMs, according to an ACL 2024 Outstanding Paper that uses the Political Compass Test as its case study. The authors, led by Paul Röttger, report that "models give substantively different answers when not forced" into fixed options, that "answers change depending on how models are forced," and that "answers lack paraphrase robustness."

The paper starts from a mismatch its authors state plainly: "real users do not typically ask LLMs survey questions." A systematic review of prior PCT work finds that most studies coerce models into the test's forced-choice format. Move the same probes into an open-ended answer setting, Röttger and coauthors write, and the models produce different answers yet again.

The framing challenges a genre of "which way does model X lean?" claims that rests on constrained benchmarks. Two experts in our Who's Who directory shared the paper on the basis of its methodology critique, not any ranked result.

Shared on Bluesky by 2 AI experts