AI Observatory: vendor usage reports filter half of chats
TL;DR
- Applied to a broader chat dataset, Anthropic's own methodology would have filtered out nearly 48% of conversations as non-work, a Stanford-led project found.
- In the excluded slice: 27.5% of chats involved harassment and hate versus 5.66% in Anthropic's analysis, and 16.7% involved sexual content versus 2.4%.
- The AI Observatory Project aggregated 85,633 conversational turns across 24,521 conversations from 5,000 users and 52 AI models between 2023 and 2025.
Applied to a broader chat dataset, Anthropic's own methodology would have filtered out nearly 48% of conversations before analysis, according to a research effort covered by MIT Technology Review and picked up by two AI experts in our Who's Who directory.
The AI Observatory Project aggregated 85,633 conversational turns across 24,521 conversations from seven datasets, involving 5,000 users and 52 different AI models between 2023 and 2025. When the group re-applied the classifier logic Anthropic uses in its Economic Index to that dataset, the 'non-work' conversations that got dropped were the ones carrying the sensitive content: 44.2% touched health and relationships versus 31.2% in Anthropic's own analysis, 27.5% involved harassment and hate versus 5.66%, 16.7% involved sexual content versus 2.4%, and 7.9% involved adult or illicit topics versus 2.1%.
The complaint is that the two data sources policymakers cite most often — Anthropic's Economic Index, drawn from about 1 million Claude conversations, and OpenAI's 2025 ChatGPT report, drawn from about 1.5 million and finding that only 30% of consumer use involved work — are self-published and unauditable.
"There is no independent source to corroborate it," says Anka Reuel, a Stanford computer science PhD candidate.
Different chatbots picked up different personas in the Observatory data. Claude concentrated in coding. Grok and Gemini leaned toward information retrieval, with Grok also concentrating misinformation and news and politics uses. Gemini drew social and roleplay uses. ChatGPT was where the homework showed up.
"No single company report tells the whole story," says Shayne Longpre of the MIT Media Lab. David Widder of the University of Texas at Austin puts the missing piece more bluntly: "When we want to ask, for example: is Anthropic's general-purpose AI system… used mostly for good or mostly for bad… we don't have a way of answering that question because that information is proprietary."
The Observatory's 24,521 conversations remain a fraction of the vendor datasets they benchmark against.
Shared on Bluesky by 2 AI experts
-
I wrote about how, beyond what AI companies tell us, we still know very about real people's AI use--and the AI Observatory project from @shaynelongpre.bsky.social @ankareuel.bsky.social + team trying to change this: www.…
View on Bluesky →
Originally reported by technologyreview.com
Read the original article →Original headline: We still don’t know how people are really using AI