arxiv.org web signal

100-Endings metric exposes LLM storytelling tension gap

TL;DR

  • Authors propose a 100-Endings metric that predicts 100 possible endings at each sentence to measure narrative tension in stories.
  • On EQ-Bench, LLM judges rank AI-generated stories above published New Yorker fiction; the new metric reverses that ordering.
  • The team builds a generation pipeline with structural scaffolding that raises measured tension while keeping EQ-Bench scores intact.

On the EQ-Bench benchmark, LLM judges rank AI-generated stories above published New Yorker fiction.

That misordering is the opening claim of "Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling," a paper by Peiqi Sui, Peter West, Ari Holtzman and co-authors. The team argues existing rubric-based evaluation systems miss narrative tension, and propose the 100-Endings metric: at each sentence of a story, a model predicts "100 possible endings based on available text." Tension is then measured, in the authors' framing, by "how often predictions fail to match the ground truth." A second geometric measure, inflection rate, tracks how often the forecast curve reverses direction, which the paper treats as a proxy for twists and revelations.

Applied back to the same corpus, the metric flips the ranking: human-written stories land substantially above LLM outputs, where rubric-based judges had put them below.

The authors then use narratological principles to build a generation pipeline with structural constraints and scaffolding that, by their account, raises measured tension while maintaining EQ-Bench performance.

Two of the researchers in our Who's Who tracker had already shared the preprint by the time it reached our scanner.

The abstract names no specific predictor model, publishes no per-genre breakdown, and gives no detail on what the pipeline's structural scaffolding actually looks like.

Shared on Bluesky by 2 AI experts