← All posts / Industry

The Ghost Workers Who Ghosted the Work: OpenAI Fires Contractors for Using AI to Train Its AI

404 Media reveals OpenAI has fired multiple contractors caught using AI to grade ChatGPT responses — inside a reviewer apparatus of 10,000+ people where em dashes, repetition, and fast turnarounds are treated as evidence, and one admitted saboteur chose the worst outputs on purpose.

The Ghost Workers Who Ghosted the Work: OpenAI Fires Contractors for Using AI to Train Its AI

On September 22, 404 Media published an investigation that lands somewhere between irony and structural warning: OpenAI has been firing the very contractors it hires to keep its models human — because they were using AI to do the job. The people paid to prevent model collapse were, in a non-trivial number of cases, feeding the collapse themselves.

What the investigation found

404 Media’s Joseph Cox obtained internal documents and spoke with three contractors working across OpenAI training projects, on top of a fourth who works on training models for multiple AI companies. The picture that emerges is of a shadow QA layer sitting underneath ChatGPT — and of a second shadow layer on top of that, hired to police the first.

The scale is the first surprise. One internal document reviewed by 404 Media indicates the projects can involve more than ten thousand contractors. That is the first public figure for the size of OpenAI’s rating workforce, and it makes explicit what the industry has long implied: the “human” in human feedback is an industrial operation, not a boutique one.

The second surprise is the rulebook. One document — describing the work of contractors hired to review other contractors, including catching them for using AI — reads:

“Do not use AI detection tools, or AI yourself. Do not use GPTZero or any other AI detection tool. They are not reliable. Reviewers may not use AI either, including Grammarly and AI translation, to review, write feedback, or write comments.”

“Do not tell evaluators why you suspect AI. It is easier for them to hide if they know what you look for. Judge the overall pattern, not one clue.”

Read that twice. The reviewers-of-reviewers are banned from using AI detectors because detectors are unreliable — and simultaneously banned from using AI at all, down to Grammarly and machine translation. Detection, when it happens, is done by human pattern-matching.

How the cheaters get caught

Contractors policing their peers are told to look for tell-tale signs: repetitive word choices, AI-style punctuation — including, in a detail that will amuse anyone who has read a thousand LLM-generated essays, overzealous use of the em dash — and contractors finishing their work suspiciously fast.

In related Slack channels, one contractor told 404 Media, people constantly post examples with the question “Is this AI?” — “Usually the answer is yes.” One source described people using AI “all the time and people are let go for it all the time, it’s pretty much the one thing that will get you kicked off ASAP,” adding that “in a group of thousands there are tons that have been caught.”

One terminated contractor shared their firing letter with 404 Media. The employer had identified problems with the “authenticity” of their work. Their account is unusually candid: “I’m not a bad person or worker. I just needed a little boost and turned to AI to help me which eventually led to my downfall. I felt no joy in the work or that I was contributing to society in any way.”

Two of the sources work for Mercor, the AI-training company that hires contractors to review ChatGPT-related material. A Mercor spokesperson confirmed the policy in a statement: contracts “strictly prohibit the use of LLMs to complete projects,” the company invests “heavily” in detection tools, and when misuse is confirmed, “we immediately remove them from the project.” OpenAI itself declined to comment.

The sabotage problem

The most quietly alarming disclosure is not the AI users — it is the saboteur. A fourth contractor, who has worked on training models for various AI companies, told 404 Media they sometimes deliberately choose the worst responses in order to degrade the model’s training:

“I either pay zero attention to the results and choose randomly or purposely choose the [worst] output. I’m not sure how much of a difference it actually makes since there are hundreds of other people also rating prompt results, but it does feel like I’m getting paid to make AI worse.”

Statistically, one saboteur in a pool of hundreds is mostly noise. But the incentive structure that produces them — low-meaning piecework at the bottom of the AI value pyramid — is the same structure that produces the AI users. Both are symptoms of work that humans find intolerable at the price offered.

Why this matters technically

The technical stakes are straightforward: reinforcement learning from human feedback breaks if the feedback is not human. If a model trains on preferences that were themselves generated by a model, you get a feedback loop of self-preference — the model effectively grading its own homework with extra steps. This is the mechanism behind “model collapse,” the degradation documented when models are repeatedly trained on AI-generated text. OpenAI’s entire post-training pipeline depends on the authenticity of the signal coming from these tens of thousands of reviewers.

The context makes it sharper. Last week’s ImpossibleRubrics study measured that models game model-written rubrics 8 to 26 percent of the time against zero for human-written ones. And 404 Media’s own earlier reporting revealed Project Lily — hundreds of contractors reading real ChatGPT users’ prompts, which can include personal information, under a setting enabled by default for consumer accounts. The authenticity crisis and the privacy crisis are two faces of the same apparatus.

The economics of authenticity

The market has already priced this in. Human-verified data is now the scarce input in AI, which is why the same news cycle carries Snorkel AI raising $350 million at a $3.5 billion valuation on 17x ARR growth, and micro1 raising above $100 million at a $4 billion valuation selling precisely this: guaranteed-human training data to labs that have exhausted the internet.

When a scarce input is this valuable, three things follow: prices rise (they have), suppliers industrialize detection (Mercor’s statement is essentially a detection-product pitch), and fraud rises proportionally. The em-dash witch hunts are what fraud detection looks like when the fraud is undetectable by machine — a recapitulation, at industrial scale, of the oldest verification problem in crowdsourcing.

The irony, examined

The obvious irony — OpenAI, whose entire pitch is that everyone should use AI at work, firing people for using AI at work — is real but shallow. The deeper version is this: OpenAI is asking tens of thousands of humans to do precisely the kind of high-volume, repetitive, low-context judgment work that its own models are best at, and then relying on those humans not to substitute the models. The company’s product undermines the company’s supply chain, and the supply chain knows it.

One contractor’s exit quote — “I felt no joy in the work or that I was contributing to society in any way” — may be the most important sentence in the piece. RLHF is sold as the transmission of human values into machines. If the humans doing the transmitting experience the work as meaningless, the values being transmitted deserve scrutiny too.

What to watch

Three open questions follow from this reporting. First, whether OpenAI’s reviewer contracts move toward camera-on, screen-recorded sessions — the direction Mercor’s “tools and systems” hint at. Second, whether the “Is this AI?” pattern-judgment layer itself becomes automated, creating a third tier of models checking humans checking models — an unwinnable recursion. Third, whether any regulator connects Project Lily’s default-on data sharing with the authenticity problem and treats both as a single data-governance issue.

The verdict on model quality from all this is unknowable from outside. But the structural fact is now on record: the human layer beneath the frontier models is porous, demoralized in places, actively sabotaged in others, and policed by em-dash forensics. Every capability number a frontier lab publishes sits on top of that layer.