arxiv.org web signal

Perspective-Driven Inference adjusts LLM labels across groups

TL;DR

  • The paper introduces Perspective-Driven Inference, treating the distribution of annotations across demographic groups as the target quantity rather than a single ground truth.
  • An adaptive sampling strategy concentrates a small human annotation budget on demographic groups where LLM proxies perform worst.
  • Tested on politeness and offensiveness rating tasks, the method improved results for harder-to-model demographic groups while maintaining coverage.

Treating LLM annotation error as a single-ground-truth problem fails for subjective tasks where demographic disagreement is itself the signal, according to a new paper by Navya Mehrotra, Adam Visokay, and Kristina Gligorić. Their method, "Perspective-Driven Inference," treats "the distribution of annotations across groups as the quantity of interest" and estimates it using a small human annotation budget.

An adaptive sampling strategy concentrates that budget on "groups where LLM proxies are least accurate," so human labor goes to correcting the model where it is worst rather than uniformly. On politeness and offensiveness rating tasks, the authors report "targeted improvements for harder-to-model demographic groups relative to uniform sampling baselines, while maintaining coverage."

The abstract names no specific LLMs, no budget sizes, and no accuracy numbers. Two researchers we follow had already shared the preprint by the time it crossed our tracker.

Shared on Bluesky by 2 AI experts