The literature on #AI "alignment" is based on utility maximization in so many papers, but most of the papers fail to even define values. Read our paper with Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, and @abeba.blacksky.app to know more. arxiv.org/abs/2608.10327
- An annotation of 94 value alignment research papers found the majority do not define values, using preferences as a stand-in.
- The authors argue this substitution risks reducing complex culturally situated concepts down to binary choices.
- A shift from human annotators to synthetic data and autoraters could close off ways to contest values in foundation models.