404media.co via Hacker News

OpenAI Fires Contractors Caught Using AI to Grade ChatGPT

5 sources tracking this story

TL;DR

  • Mercor, not OpenAI directly, executes contractor removals and confirmed an immediate-removal policy to Gadget Review.
  • OpenAI instructs supervisors to distrust automated AI-detection tools and instead rely on human pattern recognition of phrasing, punctuation, and completion speed.
  • The Curse of Recursion is the technical forcing function: repeated AI-on-AI training degrades model accuracy and output diversity over time.

OpenAI has fired multiple contractors hired to grade ChatGPT's answers after catching them using AI to do the work, 404 Media reported. Internal documents reviewed by the outlet reference "more than ten thousand contractors" across the company's rating pipelines and bar them from touching AI-detection tools, GPTZero, Grammarly, or a chatbot itself while working.

The instruction to reviewers is blunt: "Do not use AI detection tools, or AI yourself." A second line tells supervisors not to reveal how they spot cheats, since "it is easier for them to hide if they know what you look for." The tells are repetitive word patterns, excessive em dashes, and unusually fast completion times.

One fired worker's termination letter flagged issues with the "authenticity" of their submissions. The contractor's own account was direct: "I just needed a little boost and turned to AI to help me." Mercor, the staffing firm subcontracting many of these raters, told 404 Media, "When we confirm an expert has used AI to complete a task, we immediately remove them," while defending its bench: "Our experts are hired for their expertise and judgement, which is essential to the ongoing advancement of AI."

Cheating is not the only problem. One contractor described a job stripped of meaning: "I felt no joy in the work or that I was contributing to society in any way." Another described active sabotage, saying they "either pay zero attention to the results and choose randomly or purposely choose the worst output." OpenAI declined to comment.

What others are reporting

Coverage cluster as of 2h after publish

  1. Gadget Review Read →

    Connects firings to the Curse of Recursion failure mode and includes a Mercor spokesperson quote confirming the immediate-removal policy.

    When we confirm an expert has used AI to complete a task, we immediately remove them from the project.
  2. Tom's Guide Read →

    Focuses on detection mechanics: internal Slack channels flagging AI-assisted work by pattern, supervisors instructed to distrust automated tools.

  3. AI Minute Daily Read →

    Puts a scale number on the sweep (10,000-plus) and centers the structural irony that AI evaluators used the very tools they were paid to detect.

    The people paid to evaluate AI output got caught using AI to do it.
  4. Hacker News Read →

    Thread draws the data-labeler vs. AI-trainer labor distinction and frames terminations as contract enforcement rather than corporate hypocrisy.

    They hired contractors on the condition that they provide human feedback without AI; those people broke the rules, so their contracts ended prematurely.

Shared on Bluesky by 4 AI experts