AI Observatory: Anthropic method filters 48% of real AI chats
TL;DR
- Applying Anthropic's methodology to an independent dataset of 85,633 conversation turns filters out 48% of real chats as non-work.
- The excluded chats show health/relationships at 44.2% vs 31.2%, sexual content at 16.7% vs 2.4%, and harassment at 27.5% vs 5.66%.
- Users pick different models for different jobs: Anthropic for coding, ChatGPT for homework, Gemini for roleplay, Grok for news and misinformation.
Applying Anthropic's own analytic methodology to an independent dataset of 85,633 conversational turns filters out nearly half the traffic (48%) as "non-work," and the excluded chats look markedly different from what vendor reports show, MIT Technology Review reported.
The finding comes from the AI Observatory, a project co-led by Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab, and Shayne Longpre, a recent PhD graduate from the MIT Media Lab. With collaborators from MIT, Stanford and the Data Provenance Initiative, they aggregated 24,521 conversations from seven real-world datasets covering roughly 5,000 users and 52 models across 2023 to 2025.
In the conversations Anthropic's method would drop, health and relationship topics rise to 44.2% (versus 31.2%), sexual content to 16.7% (versus 2.4%), and harassment and hate to 27.5% (versus 5.66%). OpenAI's own 2025 report found only 30% of consumer ChatGPT use was work-related, suggesting the same asymmetry runs through other vendors' framings.
Different chatbots pull different behavior. The researchers found that "people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance." Grok was popular for news and politics "but it was also where misinformation tended to concentrate."
Reuel puts the underlying issue plainly: "When we want to ask, for example: is Anthropic's general-purpose AI system … used mostly for good or mostly for bad … we don't have a way of answering that question because that information is proprietary." Longpre adds that "No single company report tells the whole story," a gap we've been tracking across more than 340 Anthropic stories in the last 90 days.
Shared on Bluesky by 1 AI expert
Originally reported by technologyreview.com
Read the original article →Original headline: MIT Tech Review: AI Observatory Study Finds Industry Reports Miss Nearly Half of How People Actually Use AI