arxiv.org web signal

Illinois ISAAC bundles 527M Reddit posts across six social groups

TL;DR

  • ISAAC compiles over 527 million English-language Reddit posts from 2007 to 2023, curated around race, sexuality, age, ability, body weight, and skin tone.
  • A multi-step, human-audited filtering pipeline kept irrelevant content below 10% overall and within each of the six social group categories.
  • Every post carries algorithmic annotations for estimated home region, moralization, sentiment, emotion, and linguistic generalization, with access via website, SQL playground, Python package, and HuggingFace.

The Illinois Social Attitudes Aggregate Corpus (ISAAC) bundles 527 million-plus English-language Reddit posts from 2007 through 2023, filtered for discussion of six social group categories: race, sexuality, age, ability, body weight, and skin tone. The team behind the arxiv preprint, Babak Hemmatian, Sarah Hadjarab, Jessica Chen, and Benedek Kurdi, says a multi-step, human-audited filter kept irrelevant content below 10% both overall and inside each category.

On top of the raw posts, every item carries algorithmic annotations: an estimated home region for the author, plus labels for moralization, sentiment, emotion, and linguistic generalization. The authors argue the corpus holds up against external reality checks, pointing to "convergent evidence linking ISAAC to macro-level societal trends, such as online search behavior, temporal spikes during major societal events (both nationally and regionally), and long-term shifts in public attitudes."

The pitch is reach plus access. ISAAC "allows investigators to perform cross-category comparisons, conduct high-precision tracking of long-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes," according to the paper, with the data served through a point-and-click website, labeler web-apps, an SQL playground, a Python package, and HuggingFace. The abstract does not publish per-label accuracy numbers for the annotations, nor the method used to geolocate users.

Shared on Bluesky by 1 AI expert

  • Ted Underwood @tedunderwood.com amplified

    @benedek.bsky.social

    📊 New resource: ISAAC, an open corpus for studying social group attitudes 💬 527M+ Reddit posts (2007–2023) 🧠 Covers race, sexuality, age, ability, and skin tone 🗺️ Validated against Gallup and Google Trends 🔍 No-code or…

    View on Bluesky →