Humans are reading ChatGPT users’ prompts to improve OpenAl’s models, and those chats can include sensitive, personal information, according to leaked internal documents and real prompts seen by 404 Media.
The news presents a major privacy risk for ChatGPT’s users, with people often using ChatGPT as a therapist, professional assistant, or digital friend, and providing it with all sorts of intimate details about their lives. The contractors don’t see ChatGPT usernames, and OpenAl says it tries to remove personal information before prompts reach the reviewers, but the company acknowledged sensitive details can still get through.
The news also dispels the misconception that these models are improving only because of OpenAl’s mass scraping of the internet, the talent of its well-paid engineering and Al teams, or the power of its newer models. An important and overlooked part are the outside contractors paid to read and review ChatGPT responses to real prompts over and over again. Anthropic confirmed to 404 Media it is also using human review to improve its models.
“No,” someone who works with the prompts said when asked if they think ChatGPT users know that humans are reading their chats. “I don’t think they would imagine some contractor somewhere […] is analyzing the conversations.”
Unpaywalled link here - https://archive.ph/98Wr5
Kinda hate to show my whole ass on this topic, but I lived through this exact thing 5 or 6 years ago when AI Dungeon accused its entire userbase of being pedophiles.
Were they?
Why is that your question?
Was the AI dungeon user base pedofiles?
I’m 100% positive they’re not reading my ChatGPT logs.
(I don’t use ChatGPT, and neither should anyone else.)
You and me… same boat
Someone can submit your messages / cv etc to chatgpt against your will.
Sure, but how likely is that?
They’ll be scraping all of Lemmy too for sure…
Oh no, I’m so flabbergasted! A US-bigtec’s chatbot’s “conversations” are monitored and read? How shocking and unexpected.
Maybe the best tactic to keep confidential information confidential is, well, to keep it confidential and not put into some chatbot. But what do I know…
There probably has been some disclaimer that we all got so used to just clicking away. Saying that everything you say here will be recorded and read by 3rd parties. Or some such. Right?!
I’m not saying this makes it OK though.
some disclaimer that we all got so used to just clicking away.
That’s what LG said… https://www.theverge.com/tech/994333/lg-responds-to-tv-spying-allegations
Mechanical Turk, redux.
https://en.wikipedia.org/wiki/Mechanical_Turk
Also, the privacy concerns are an absolute nightmare. 100% chance this leads to a massive leak.
That’s not what this is. Humans aren’t providing the responses, they’re reviewing them, grading them, and using them to train models.
An example of a mechanical turk would be Waymo using humans to navigate situations the computer can’t.
If AI is so smart, why does it need this level of baby-sitting from humans? Mechanical Turk, one step removed.
The tech gods are saying they’ve reached Artificial General Intelligence, yet clearly if they have to pay humans to continually train the bots on the subtleties of human interaction, they have not reached AGI, and even feeding an LLM the entire corpus of the internet is insufficient to train them.
As much as I think this is a good and necessary step (privacy concerns aside), it definitely does point to AI being significantly less than advertised.
It’s not babysitting, it’s training. The humans grading the responses is how the models are being improved. It’s just a more targeted approach than “let’s scan the whole internet.”
Humans aren’t pretending to be AI. This is a completely different scenario than the mechanical turk.
You are correct, it is not literally talking to a person. You could even go further and point out that the trainers are not literally Turkish.
You seem to not be grasping the concept of the mechanical turk. The human was the one playing chess, not the machine. The machine playing chess was a lie.
You’re basically saying that a student’s math teacher is actually doing the student’s homework because they taught the student how to do it.
Yes, and this machine that taught itself how to be smart from absorbing data is likewise a lie.






