If you're at #COLM2026, come check out our poster for "Spoiler Alert" (arxiv.org/abs/2604.09854). We ask why LLM fiction is so bad at holding narrative tension, and create a metric to measure tension in short stories. What improves LLM fiction on this metric, it turns out, is …
- Authors propose a 100-Endings metric that predicts 100 possible endings at each sentence to measure narrative tension in stories.
- On EQ-Bench, LLM judges rank AI-generated stories above published New Yorker fiction; the new metric reverses that ordering.
- The team builds a generation pipeline with structural scaffolding that raises measured tension while keeping EQ-Bench scores intact.