docs.google.com web signal

Princeton lab urges psychologists to avoid proprietary LLMs

TL;DR

  • M.J. Crockett of Princeton's psychology department publicly recommends avoiding proprietary LLMs like GPT, Gemini and Claude in scientific research.
  • The resource cites Barrie et al. (2024): retired proprietary LLMs cannot be replicated, and outputs can drift over time even for identical model-and-prompt pairs.
  • It hands reviewers a copy-pasteable paragraph challenging papers that use proprietary LLMs on transparency and reproducibility grounds.

M.J. Crockett, a Princeton psychology researcher, is telling her field to stop using OpenAI's GPT, Google's Gemini, or Anthropic's Claude in scientific research. Her public resource, last updated 10/2/25, catalogues what she calls the "ethical costs and epistemic risks" of these models and hands colleagues a copy-pasteable peer-review paragraph to send back to authors who rely on them.

The recommendation is unambiguous: "avoid using proprietary LLMs in scientific research, unless your research question is specifically about understanding the characteristics of proprietary LLMs." Crockett borrows the term "proprietary LLMs" from Barrie et al. (2024), meaning models whose weights, source code and training data are kept hidden as a "trade secret," and names GPT, Gemini and Claude as her examples.

The reproducibility argument leans on Barrie: "Proprietary LLMs are regularly retired; once they cease to exist, it's not possible to replicate their outcomes." And: "as researchers update to newer and newer models… there is no reason for them to expect that effect sizes from earlier models will be preserved." Even the same model and the same prompt, she writes, can produce different outputs over time in ways that affect research conclusions.

The ethical section is more sweeping. It links to reporting that "OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic," to copyright suits filed against OpenAI, Microsoft, Anthropic and Google, and to Bender et al. (2021)'s claim that "LLMs' training data overrepresent the voices of people most likely to adhere to hegemonic perspectives." Two of that paper's authors, Gebru and Mitchell, were "forced out of Google shortly after the paper was published," Crockett writes.

The document is a curated reading list, not a peer-reviewed study, and it flags that itself: "It is not meant to be comprehensive; it is a curated list of readings that shaped my personal views on how to best use LLMs." Two of the researchers we track shared its link.

Shared on Bluesky by 2 AI experts