↻
David Smith reposted
Ted Underwood
@tedunderwood.com
Paper: arxiv.org/abs/2609.23178 Code: github.com/Historical-A... Team: Ted Underwood, Ziliang Qiu, Laura K. Nelson @lauraknelson.bsky.social, Sarah Griebel @sgriebel.bsky.social, Edwin Roland @teddyroland.bsky.social, Wenyi Shang @wenyishang.bsky.social , Matthew Wilkens @matt…
AI Weekly's analysis
→
- The Chronologic benchmark evaluates language models on English-language contexts from 1831 to 1930 using pairwise comparisons and strong distractors.
- Generative tasks proved harder than discriminative ones, and reasoning models could typically identify weaknesses in their own generated answers.
- Models pretrained only on historical text led on answer likelihood but could not match commercial models in free generation.
Read full analysis →
View on Bluesky →