link.springer.com web signal

Sociologists map how LLMs are reshaping text-as-data research

TL;DR

  • Four researchers from Linköping, Leipzig, Utrecht and Uppsala organize computational text analysis into data-first, theory-first, and theory–data integration paradigms.
  • Foundation LLMs like GPT, Llama and Claude can now be deployed across all three paradigms, softening earlier boundaries between unsupervised, supervised and semi-supervised methods.
  • The authors flag hallucinations, prompt sensitivity, proprietary-API data risks, and SUTVA violations as unresolved barriers to defensible causal inference from text.

A new open-access piece in the Kölner Zeitschrift für Soziologie und Sozialpsychologie tries to give sociology a shared vocabulary for what has become a very crowded methodological corner. Four researchers based at Linköping, Leipzig, Utrecht and Uppsala argue that computational text analysis has matured from a descriptive novelty into a measurement toolkit for building and testing social theory, and they lay out how the pieces fit together.

The organizing move in the paper is a three-way split. Data-first methods, such as topic models, word embeddings and BERTopic, discover latent patterns that researchers interpret afterward. Theory-first methods, including dictionaries, supervised classifiers, and few- or zero-shot classification with models like GPT, require the researcher to specify categories in advance. Theory-data integration methods, such as structural topic models, seeded topic models and interpretative word embeddings, sit between the two. The rise of foundation large language models has begun to soften these boundaries, because the same GPT, Llama or Claude model can be pointed at any of the three tasks depending on how it is prompted.

What the authors are careful not to oversell is causal inference. Most text-as-data work, they write, has remained primarily descriptive, and moving to causal claims runs into a technical problem that Egami and colleagues named in 2022: the way a researcher builds a codebook on one slice of documents shapes how every other document gets labelled, violating the stable-unit-treatment-value assumption that standard causal designs rely on. The recommended fix is data splitting, and the authors concede that few current studies, including their own, implement it.

The honest caveats sit in the discussion. Generative LLMs deliver strong headline performance on some annotation tasks, but the reporting also flags inconsistent outputs across runs, model hallucinations, sensitivity to subtle variations in prompting, biased representations of minority social groups and languages, and the reproducibility problem of sending research data through proprietary APIs. On the newer fashion of using LLMs to simulate survey respondents, the authors are pointed: current models tend to underestimate the diversity of responses given by actual humans, and skew toward majority opinions.

What the reporting does not give you is a benchmark table telling you which model to use for which task, or the cost and licensing tradeoffs. What it does give you is a framework, and, for teams already piping corpora through Claude or GPT for research they intend to publish, a checklist of the failure modes worth validating against before the paper goes out.

Shared on Bluesky by 1 AI expert

  • Ted Underwood @tedunderwood.com amplified

    @mhbodell.bsky.social

    How can computational text analysis be used for building and testing social theory? @marckeuschnigg.bsky.social, @anamacanovic.bsky.social, @anmen.bsky.social & I write about how CTA is commonly used & how it might be …

    View on Bluesky →