nature.com web signal

Pangram catches AI writing at scale, but edits still slip

TL;DR

  • NeurIPS rejected 18% of submissions to one of its tracks in June after screening them with the AI detector Pangram.
  • Pangram's technical paper reports a 0.0041% false-positive rate on English texts, with similar rates across more than 100 languages.
  • But Pangram still labels heavily AI-edited human student essays as fully human 41% of the time, per its own tests.

Nature reports that a new class of AI text detectors, led by Pangram, has become accurate enough for major publishers and conferences to rely on, though the tools still miss essays that were only lightly rewritten by AI.

In June, NeurIPS said it rejected 18% of submissions to one of its tracks after screening them with Pangram. The American Association for Cancer Research found that 23% of abstracts in manuscripts and 5% of peer-review reports submitted to its journals in 2024 contained text probably generated by large language models, and fewer than 25% of authors disclosed any AI use despite the publisher requiring it. AACR now runs Pangram over every incoming peer-review comment. A separate study reported in January found that one in eight biomedical articles last year contained some AI-generated text.

Pangram's technical paper reports a 0.0041% false-positive rate on English texts, with similarly low rates in more than 100 languages, and, unlike earlier detectors, no bias against writers who are not fluent English speakers. The same tests, though, found that when consumer AI was used to "substantially modify" human-written student essays, Pangram still labelled the results fully human 41% of the time. Wholesale generation gets caught. Heavy assistance often does not.

The company's co-founder Max Spero, quoted in the piece, frames the stakes as cultural: "If it's taboo to call out that somebody's using AI to write, then I think we're going to see a lot more people shirking their jobs and letting AI replace themselves. We're in a really critical time of setting norms." Competitor GPTZero, based in New York City, markets itself as offering 99% accuracy and "the most precise, reliable AI detection results on the market."

Shared on Bluesky by 4 AI experts