Paper: sycophancy evals conflate receptiveness with deference
TL;DR
- A new paper argues that markers of 'social sycophancy' in language models overlap with conversational receptiveness, penalizing desirable behavior.
- In a preregistered experiment on moral-advice responses, more receptive answers were preferred even by participants who disagreed with the asker.
- The authors describe a simple approach that raises receptiveness in model outputs without increasing substantive deference.
Sycophancy evals may be measuring the wrong thing. A new preprint on arXiv, posted September 22 by Calvin Isley, Johann Gaebler, Max Lamparth, Julia Minson and Sharad Goel, argues that behaviors flagged as 'social sycophancy' in language models overlap with something social psychologists call conversational receptiveness: the warmth and openness that make disagreement productive.
'A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment,' the authors write. But when they took human-written moral-advice responses and increased the receptiveness of the language while keeping the substantive conclusion unchanged, those responses got classified as more sycophantic. The tone moves, the position does not, and existing metrics fire.
That is the construct-validity problem the paper is pointing at. In a preregistered experiment, 'participants prefer the more receptive responses, expect users to be more likely to listen to them, and are more willing to seek advice from their authors.' The preference held up even among participants who disagreed with the original question asker.
The authors say 'conversational receptiveness and substantive independence can be achieved together,' and describe a simple approach that raises receptiveness in model outputs without adding deference. Two researchers on our tracker had shared the link by the time we picked it up.
Not visible in the abstract we retrieved: which specific language models were tested, the sample size of the human experiment, or the size of the preference gap.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models