Chronologic benchmarks LLMs on 1831-1930 English contexts
TL;DR
- The Chronologic benchmark evaluates language models on English-language contexts from 1831 to 1930 using pairwise comparisons and strong distractors.
- Generative tasks proved harder than discriminative ones, and reasoning models could typically identify weaknesses in their own generated answers.
- Models pretrained only on historical text led on answer likelihood but could not match commercial models in free generation.
Not one of the language models tested can persuasively represent English-language contexts from 1831-1930, though 'progress toward that goal is evident,' the authors write. The benchmark, called Chronologic, is posted to arXiv by a team led by Ted Underwood.
The team scored models with pairwise comparisons to multiple ground truths and strong distractors. Discriminative tasks were easier than generative ones. Reasoning models, the paper says, 'can typically discern the weakness of their own generated answers.'
The interesting split is by training diet. Models pretrained exclusively on historical text led when evaluated by answer likelihood but 'cannot compete with commercial models in free generation.'
Shared on Bluesky by 1 AI expert
Originally reported by arxiv.org
Read the original article →Original headline: Chronologic: Measuring Language Models' Ability to Represent the Past