Smart, Birhane audit finds AI alignment rarely defines values
TL;DR
- An annotation of 94 value alignment research papers found the majority do not define values, using preferences as a stand-in.
- The authors argue this substitution risks reducing complex culturally situated concepts down to binary choices.
- A shift from human annotators to synthetic data and autoraters could close off ways to contest values in foundation models.
A new arXiv paper from Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, Kris Shrishak and Abeba Birhane goes back to the basics and asks a question the alignment literature mostly skips: what does the field actually mean by 'human values'? Their answer, after annotating 94 value alignment research papers, is that most of the papers do not define values at all. What they use instead is preferences as a stand-in.
Preferences are convenient because they are cheap to collect at scale and slot neatly into optimization. The trade-off the paper flags is that this substitution 'runs the risk of reducing complex culturally situated concepts down to binary choices'. If your training signal is a thumbs-up on option A over option B, whatever is thick and contested about a value has already been flattened into a preference by the time the model sees it. Two tracked researchers in AI Weekly's Who's Who directory circulated the arXiv link, a small signal the alignment community is treating this piece as worth reading.
The second concern is where the pipeline is going. As researchers 'dispense with using human annotators for model training and evaluation, turning instead to synthetic data and autorater approaches', the authors warn this could 'close off alternative methods for contesting and enacting values in foundation models'. The framing here is about process, not just output: strip out the humans and you strip out one of the few places where a contested value can actually get contested.
The paper is a philosophical audit rather than a benchmark result, so a few gaps sit alongside the argument. It does not tell you which specific companies or model families are among the 94 papers, nor does it propose a replacement operational definition of 'values' that alignment teams could plug in tomorrow. The point is diagnostic: making the field's implicit commitments explicit so a fight about what 'alignment' even means can happen in the open.
For anyone building or buying 'aligned' models, the practical takeaway is skepticism about what the label promises. If your vendor's alignment story reduces to preference optimization and autorater scoring, that is a specific, contestable choice with cultural weight, not a neutral technical fact.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Toward a Theory of Value in AI Alignment