OpenAI's 'Project Lily' Pays Contractors to Read ChatGPT Chats
TL;DR
- OpenAI's internal 'Project Lily' pays hundreds of outside contractors over $50 an hour to read and rate real ChatGPT conversations.
- Reviewers score four candidate responses on a 1-7 scale, judging tone and behavior rather than factual accuracy.
- Free, Plus, and Pro accounts have model training on by default; Enterprise, Business, and Education tiers do not.
OpenAI is paying hundreds of outside contractors over $50 an hour to read and rate real ChatGPT conversations under an internal program called Project Lily, according to Joseph Cox's September 14 investigation for 404 Media. Reviewers score up to four candidate responses on a 1-7 scale.
They are not grading facts. They are grading manners: whether the bot sounds too much like a bot, whether it claims to be human, whether it eases off the reflexive agreement that lawsuits have tied to the "over-sycophantic 4o model led in part to multiple peoples' suicides."
OpenAI strips account names and runs prompts through automated filters before transcripts reach reviewers. The filters miss things. Short exchanges are especially leaky, and the "user memories summaries" the model builds up routinely surface a person's job, geographic location, and background alongside the chat.
"I don't think they would imagine some contractor somewhere [..] is analyzing the conversations," one worker told 404 Media, referring to the 900+ million people using the product. Free, Plus, and Pro accounts have model training on by default; Enterprise, Business, and Education accounts do not. Anthropic confirmed it runs a similar review process.
Two of the AI analysts we follow in our Who's Who directory circulated the podcast the day it dropped.
Shared on Bluesky by 2 AI experts
Originally reported by youtube.com
Read the original article →Original headline: Humans Are Reading Your ChatGPT Conversations