arxiv.org web signal

OLMo3-7B paper splits social and STEM reasoning by data source

TL;DR

  • OLMo3-7B's social and STEM reasoning trace to 'qualitatively distinct corpus regions' of its Dolma3 pretraining mix, per gradient-based training-data attribution.
  • The social-STEM split is sharper at the reasoning level (SocialIQA, ARC-Challenge) than at the knowledge level (MMLU Social Sciences, MMLU STEM).
  • Unlearning high-attribution topics like Literature degraded SocialIQA more than within-topic random baselines, giving what the authors call 'partial causal validation.'

OLMo3-7B's social reasoning and its STEM reasoning draw on "qualitatively distinct corpus regions" of the model's pretraining data, according to a 101-page paper by Glenn Matlin, Stella Biderman, Mark Riedl and colleagues, posted to arXiv and published at COLM 2026.

The authors ran gradient-based training-data attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, then aggregated the per-document influence scores into WebOrganizer's 24-format by 24-topic taxonomy, yielding 576 bins. Four benchmarks were arranged in a 2x2 by domain and capability type: SocialIQA and MMLU Social Sciences on the social side, ARC-Challenge and MMLU STEM on the other. The contrast between social and STEM, the paper reports, is "sharper at the reasoning level than at the knowledge level."

Causality is tested with targeted machine unlearning: "forgetting high-attribution topics (e.g., Literature for SocialIQA) degrades the aligned benchmark more than within-topic random baselines." The authors frame this only as "partial causal validation," not proof.

Code, data artifacts, influence scores and checkpoints are released on GitHub and Hugging Face. Three researchers we follow in our Who's Who directory shared the arXiv link.

Shared on Bluesky by 3 AI experts