TrueSeeker AI · Verified claim report Case 4752afe45e · 2026-09-18

§ Claim under review · Business

"OpenAI is reportedly paying human contractors to read real ChatGPT users' chats, under an internal project codenamed 'Project Lily', according to 404 Media."

Circulating claim, as submitted.

Verdict

Mostly accurate

Confidence

High
§

Summary

This one checks out in its core. 404 Media published an investigation on 14 September 2026, based on leaked internal documents, reporting that OpenAI pays hundreds of contractors to read real ChatGPT conversations and rate the chatbot's answers, under the internal codename Project Lily. OpenAI did not deny it, and told reporters that expert raters may review users' conversations, and OpenAI's own help pages confirm that data sharing is switched on by default for Free, Plus and Pro accounts and that opting out applies to new conversations. The post's weak point is what it leaves out: it repeats OpenAI's assurance that usernames are stripped and a privacy filter redacts personal details, but omits that OpenAI itself says that filter can miss things, that reviewers can see a summary of a user's memories including approximate location, and that opting out does not undo past eligibility. The post also omits that Anthropic confirmed it uses human review too and Google discloses the same, so this is an industry practice rather than something unique to OpenAI. What remains unconfirmed is anything resting only on documents no one else has seen, including the codename itself, the contractor headcount and which model is being trained. Anyone treating ChatGPT as a private diary should know that human review of consumer chats is the documented default unless the setting is turned off.

§

The readings

key figures from the evidence
50 USD/hour

contractor pay rate for reviewing ChatGPT conversations

1 to 7

rating scale contractors use to score chatbot replies

hundreds contractors

contractors reportedly hired to manually review ChatGPT logs

§

Why this verdict

The named source is real, the attribution is exact, the codename is quoted correctly, and the post hedges appropriately with "reportedly." Unusually for a single-outlet document leak, the core proposition is corroborated by the subject itself: OpenAI told a second outlet on the record that expert raters may review users' conversations, and OpenAI's own help pages confirm default-on consumer data sharing, the opt-out, and the Privacy Filter's documented fallibility. I considered and rejected "Credibly reported but unconfirmed," because that verdict requires no on-the-record confirmation by any party and here OpenAI confirmed the practice while leaving only the codename and headcount unconfirmed. I rejected "Accurate" because the caption reproduces OpenAI's reassurances while dropping the acknowledged limits that sit beside them in the source, and I rejected "Partially accurate but misleading" because those omissions soften the risk framing without altering the operative proposition, which stands fully supported as of 2026-09-18.
§

Evidence

The named source exists and says what the post says it says. 404 Media reports that humans are reading ChatGPT users' prompts to improve OpenAI's models, and that those chats can include sensitive personal information, based on leaked internal documents and real prompts seen by the outlet, with OpenAI hiring hundreds of contractors who read a stream of real users' ChatGPT prompts

The stated goal of these prompt review teams is to improve the responses ChatGPT gives to users, with contractors rating and critiquing the chatbot's generated replies . Internal documents seen by the outlet show contractors training ChatGPT not to anthropomorphize itself and to be less sycophantic, described as a key problem for OpenAI .

On the privacy mechanics the post cites: the contractors do not see ChatGPT usernames, and OpenAI says it tries to remove personal information before prompts reach reviewers, but the company acknowledged sensitive details can still get through . Reporting on the same story states OpenAI scrubs usernames and routes text through an automated tool called Privacy Filter before conversations reach reviewers, that OpenAI acknowledges the filter can fail on rare identifiers or ambiguous phrasing, and that contractors can still view a "user memories summary" showing a person's past interests and approximate location .

The Privacy Filter is a real, documented OpenAI artifact rather than a claim invented for this story. OpenAI's own documentation states that the filter should be used as part of a privacy-by-design approach and not as a blanket anonymization claim , and advises keeping human review paths for high-sensitivity workflows .

On the opt-out element, OpenAI's own help pages corroborate the post and add the two qualifiers the post leaves out. OpenAI states that on ChatGPT Plus, Pro or Free personal workspaces data sharing is enabled by default, that users can opt out, and that once a user opts out new conversations will not be used to train its models .

On mechanics and pay, reported by the single originating outlet: contractors are described as earning over $50 an hour through intermediary firms Crossing Hurdles and Mercor, reviewing selected conversation streams and scoring generated replies from 1 to 7 .

On subject response: OpenAI did not deny the story. Asked by a separate outlet whether users were made explicitly aware their prompts may be used for training or refinement, OpenAI did not directly address the question, and a spokesperson instead emailed a bulleted list stating the company makes clear that human feedback helps improve response quality and that expert raters may review users' conversations and assess model responses, along with a link to its "data usage for consumer services" page, which 404 Media says OpenAI updated after being contacted for comment . Tom's Hardware reports that OpenAI initially had no answer on whether users were explicitly told chats could be read by humans and later pointed to an FAQ page on human review, which that outlet verified is at least two years old .

Context the post omits: the same reporting notes that an overlooked part of model improvement is outside contractors paid to read and review responses to real prompts, and that Anthropic confirmed to 404 Media it is also using human review to improve its models . Coverage also notes Google's Gemini states in its Privacy Center that some saved chats may be reviewed by humans, and that Anthropic holds a similar position with a dedicated page explaining it .

§

Findings

✓ What's accurate 6

  • The 404 Media article exists, is attributed correctly, is dated 2026-09-14, and is by Joseph Cox. The post's "reportedly" and "according to 404 Media" hedges are accurate attribution, not laundering.
  • The codename "Project Lily" appears in the reporting as the internal codename found in leaked documents, exactly as the post states.
  • The described workflow matches the source: reviewers read real user prompts, summarize intent, compare multiple candidate responses, and rate them.
  • The anthropomorphizing and sycophancy element matches the reporting, including the link to the GPT-4o sycophancy problem.
  • The two OpenAI statements the post reports are accurately reported: usernames are not shown to reviewers, and a privacy filter is applied before review.
  • The opt-out statement is corroborated by OpenAI's own help documentation, not just by the report.

≈ What's misleading 5

  • Omitted qualifier: the post relays OpenAI's privacy-filter reassurance but omits the acknowledgment sitting next to it in the same reporting, that the filter can miss uncommon or ambiguous identifiers and that sensitive details can still reach reviewers. OpenAI's own model documentation says the filter is not a blanket anonymization claim. Presenting the safeguard without its stated failure mode converts the investigation's central privacy finding into a reassurance.
  • Omitted qualifier: the post says chats are not used for model improvement when users turn off "Improve the model for everyone," without noting that the setting is on by default for Free, Plus and Pro accounts, and that OpenAI's own page frames the effect as applying to new conversations. A reader could infer an opt-in regime and a retroactive withdrawal, and neither is what the documentation describes.
  • Omitted qualifier: the post does not mention the "user memories summary" visible to reviewers, which the reporting says can surface past interests and approximate location. This is the specific mechanism that undercuts the "usernames are removed" reassurance the post does include.
  • Omitted context, not a distortion of the claim itself: the reporting states Anthropic confirmed it also uses human review and that Google discloses the same practice. The post's framing leaves the practice looking OpenAI-specific rather than industry-standard.
  • Minor imprecision: "OpenAI is paying human contractors" compresses a chain in which recruitment runs through Crossing Hurdles and payment runs through Mercor. This is a reasonable simplification and does not change the operative proposition.

? What's uncertain 5

  • The leaked internal documents are not public. Every finding that rests on them, including the codename "Project Lily," the headcount, the 1 to 7 scale and the pay rate, rests on one outlet's description of material only it has seen. OpenAI has not confirmed or denied the codename.
  • I retrieved only the ungated portion of the 404 Media article. Details inside the member-gated remainder are known to me through secondary coverage rather than direct retrieval.
  • Which model or models Project Lily is training is not established; coverage explicitly notes the materials do not identify it.
  • How prompts are selected for review, and how many, is not established.
  • Whether OpenAI's pre-story disclosures were adequate is contested rather than resolved. OpenAI points to existing help language; reporters say the company would not directly answer where users were told, that the cited FAQ predates the story by years, and that a help page was edited after the outlet made contact. I did not independently retrieve archived versions to verify the edit timeline.
Distortion flags omitted qualifier
§

Sources

7 of 9 linked to records
[1]

404 Media, Joseph Cox, "Inside 'Project Lily': The Humans Reading Your ChatGPT Chats," published 2026-09-14

primary named-outlet investigative journalism, worker-owned, editorially accountable
https://www.404media.co/inside-project-lily-the-humans-reading-your-chatgpt-chats/ ↗
[2]

OpenAI Help Center, "What if I want to keep my history on but disable model training?"

primary vendor channel of record
https://help.openai.com/en/articles/8983130-what-if-i-want-to-keep-my-history-on-but-disable-model-training ↗
[3]

OpenAI Help Center, "How your data is used to improve model performance" and "Data Usage for Consumer Services FAQ"

primary vendor channel of record
https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance ↗
[4]

OpenAI Privacy Filter model card and repository

primary vendor documentation
https://github.com/openai/privacy-filter ↗
[5]

Yahoo Tech, "Humans may be reading your ChatGPT conversations. Here's how to opt out"

secondary named-outlet journalism
https://tech.yahoo.com/ai/chatgpt/article/humans-may-be-reading-your-chatgpt-conversations-heres-how-to-opt-out-152059277.html ↗
[7]

TechRepublic, The Next Web, Futurism, IBTimes UK

secondary tech press, all attributing to 404 Media
This citation could not be independently verified.
[8]

Joseph Cox on X announcing the story

secondary author self-promotion, confirms authorship and framing
https://x.com/josephfcox/status/2099505045481906224 ↗
[9]

ForkLog, Gigazine, LatestLY, Gadget Review, KuCoin, LavX, Tom's Guide, GB News, aiweekly.co, a Penn State blog repost

tertiary aggregation and syndication, zero independent evidentiary value
This citation could not be independently verified.
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →