arxiv.org web signal

AI graders penalize 'AI-generated' prose 2.5x more than humans

TL;DR

  • Thirteen AI evaluators graded passages labeled 'AI-generated' 34.3 percentage points lower than identical passages labeled human, versus 13.7pp for 556 human readers.
  • The bias held across a 14x14 matrix of AI creators and evaluators at +25.8pp, indicating it operates across model architectures.
  • Attribution labels caused evaluators to 'invert assessment criteria,' rating identical features positively or negatively based solely on perceived authorship.

Show a large language model a paragraph of literary prose, tell it a human wrote it, and it will grade higher than an identical paragraph labeled "AI-generated." 556 human readers do the same, just much less strongly. In a controlled study using Raymond Queneau's Exercises in Style, humans showed +13.7 percentage points of preference for passages labeled human-authored. Across 13 AI evaluators, the effect ran to +34.3 percentage points, "a 2.5-fold stronger effect (P<0.001)."

The setup, described in a preprint on arXiv by Wouter Haverals and Meredith Martin, pitted Queneau's originals against GPT-4 rewrites under three conditions: blind, accurately labeled, and counterfactually labeled. A second study ran a 14x14 matrix of AI evaluators against AI creators to see whether the pattern held across architectures. It did, at "+25.8pp."

The authors describe not a small tilt but a full inversion. "Attribution labels lead evaluators to invert assessment criteria," the paper reports, "with identical features receiving opposing evaluations based solely on perceived authorship." Their read is that models absorbed the human cultural bias against machine creativity through training and alignment: "AI systems not only replicate but amplify this human tendency."

Two researchers in our expert tracker posted the link this month.

Which of the 13 models led the effect, and whether alignment-heavy or reasoning-tuned models performed differently from base instruct models, is not stated in the abstract. The specific training-pipeline source of the amplification is also not identified.

Shared on Bluesky by 2 AI experts