- cross-posted to:
- [email protected]
- cross-posted to:
- [email protected]
Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”
Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.
It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.
What seriously worries me about these companies is the fact they somehow act like what ChatGPR is doing is 100% out of their control. Dude, you built the LLM, you trained it on stolen data.
Stop acting like “oh evil LLM is built on theft and is hacking all these companies and will doom us all, whatever will we doooo?!” when it was your damn company that started this whole crap and you’re always blabbering about general AI and how this time it will definitely arrive by 2029 if they give you another 10 billion dollars. I can’t with these clowns. Remember when Musk had us believe he wanted to save the world and came from poverty? Yeah, good times
weird how Chinese AI firms aren’t saying this - but then, they don’t have a massive fiscal crunch where they have to take delivery of trillions of dollars worth of contracted compute, that they haven’t yet found a market for . . . .
https://www.groundbrkr.com/p/the-teaser-period-why-the-ai-boom
If I didn’t know better, I’d think they were trying to angle for a bailout!
Relax guys, the very stable high IQ president of the united states will protect you. No meed for guardrails or regulations.
“Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said.
Microsoft executives, including CEO Satya Nadella, testified under oath that after ripping content from the New York Times and other news sites, clicks to those news sites fully cratered, falling by more than 90 percent on Bing.
Documents obtained during the court proceedings found that OpenAI created “a hack to get around nytimes paywall,” to which OpenAI cofounder Greg Brockman said “ah, nice.” Microsoft executive Brent Hecht wrote that LLMs steal content “without ways of distributing economic value down the supply chain, [which] necessarily threatens the economic stability of those who create the content.”
Fuck these guys.

We’ve known since the late 1800’s that the end game of capitalism is to automate all labor.
someone should write a book about how capitalism functions. Call it like The Capital or something.
That’s the goal of most technology. To allow more things to get done in less time with fewer employees.
CNC machine, tractor, chainsaw, printing press for some examples.
That’s because, thus far, they get away with choosing not to distribute any of their trillions of dollars to the suppliers of the information they consume - money has only gone to the suppliers of hardware and power, and to influencing politicians and rewarding investors. That’s their choice, and they should not be allowed to continue to make that choice. Good luck to the NYT.
https://www.advocate.com/politics/national/new-york-times-transgender-controversy
https://www.npr.org/2023/02/15/1157181127/nyt-letter-trans
https://theintercept.com/2024/04/15/nyt-israel-gaza-genocide-palestine-coverage/
NYT is transphobic and pro-genocide. I hope they go bankrupt, and papers with fewer conservative biases survive. Unfortunately, the opposite will happen. Billionaire-backed news will run at a loss to manufacture propaganda, while independent journalism dies.
True enough. Maybe in an ideal world, they would both lose.
I honestly don’t see a point where we can move past this without Sam altman, nadell, musk, etc all facing mandatory life time sentence without parole, work release etc.
I would also accept the removal of their heads.
I’d prefer if they were forced to pay royalties for all the work they stole to the people they stole it from. And, like, have someone actually force them to comply. They’d have to hire a Microsoft-sized compliance department just to figure this shit out and track payments.
Yeah, but with which money? All those AI companies are not and will never be profitable (and they’ll be the reason for the next stock market bubble pop)
Which is exactly why they are pivoting to “AI sucks, LLMs are dangerous…”
Better yet, seize their assets and make them try to make a go at it as a regular person.
Seize their assets and use that to pay creators whose works they stole.
I like how you think outside of the box.
Could you imagine Elon Musk running a fryer or injection mold for 8 hours? Who am I kidding? he’d never pass the piss test required to run the injection mold.
The best succinct description of llm that I’ve read was “plagiarism machine”, because that’s basically what they are: automated plagiarism. At best they create a collage of other works, at worst they copy one work verbatim, but they never create something original because they can’t.
Copying something is a lot easier and thus cheaper than creating an original work from scratch, so genuine creators cannot hope to compete on cost. Which leads to less people being able to afford to earn a living from creating works, which leads to less original works being created.
Short term that’s great for the neoliberal company executives: fire the expensive artists, create cheap slop with the plagiarism machine, and thus maximize profits now. That the plagiarism machine becomes stagnant because not enough new original works are being created is a future problem, by which time the current executives will have already jumped ship.
But while it’s great for the neoliberal business model, for the rest of society it will suck. So now that we have confirmation that the AI executives know how bad their products are, and that they were just publicly lying about it in the style of tobacco executives, will anything be done about this “astonishing theft of unprecedented proportions”? Personally I doubt it, there’s too much regulatory capture.
don’t base your thoughts on the plagiarism machine that can’t create anything new because that idea isn’t correct
Internet search is terrible now. Every website I go to reads like it was generated by an LLM.
Because it was generated by a llm.
I did a search for the torque spec on a castle nut last week and the second result was for a nut (as in, food nuts that you eat) website and the llm had gone all in on nuts and added a page about castle nuts. It even generated an image of a castle for the top because… castle nut.
It proceeded to provide vague instructions and a torque recommendation of 2x the actual spec.
Then there was the other one about a small 4 stroke engine where it stated “other sites will say you don’t need to mix oil with gas but you absolutely must!” (You don’t and shouldn’t) … I wonder how many people have fucked up their shit by trusting llm garbage without knowing better.
Running into this same problem but with parenting. My wife is constantly asking me to “Google if it’s good/bad if our baby is doing” X. And I’ll find 20 slopposts from dozens of “parenting” websites that all give conflicting answers. This week, I asked my parents for their copy of Dr. Spock’s Baby and Childcare because we need some solid reference. Sorry kiddo, you’re going to get raised like it’s 1987 because 2026 sucks.
Just give the baby a cigarette. It lubricates the lungs.
Plus they’ll look sick af
They knew what they were doing every step of the way. They are criminals that need to be disarmed and locked away. And we need to create a new Internet from scratch somehow thanks to these donkeys.
If we make a new internet from scratch can we get rid of JavaScript too while we’re at it?
Here you go:
https://en.wikipedia.org/wiki/Gemini_(protocol)
https://github.com/kr1sp1n/awesome-gemini
(This protocol does not have to do anything with Google’s Gemini thing. It was created and named a bit earlier!)
Microsoft, in a policy document, wrote that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained […] LLMs are a product that destroys its supply chain.”
If that’s not the very definition of a categorically unsustainable business model, I don’t know what the fuck is.
They don’t care. They want all the money this quarter.
a “doom loop” that is killing “the entire web.”
the corporate web maybe. how does it affect independent parts like the fediverse?
Something others haven’t touched on is the burden all the scrapers are putting on small websites and forums.
Lots of niche/hobbyist site runners have repeatedly talked about how the scrapers inefficiently scrape the same pages, images, and files endlessly. Until they crash the site, run them out of bandwidth, or run up their bills till they cant afford to host a site they’ve been running for many years before this nonsense.
All in the futile pursuit of more to dump into the machine because of the belief if they just had more data all the issues with the models would finally be solved.
Slop commenting. The litany of AI slop websites that have popped up, polluting search results. The absolute flood of AI slop pull requests being submitted to open source projects, to the point at which many open source maintainers are buckling under the weight of garbage being thrown their way. Bots everywhere. The list goes on and on and on.
The internet is already a dessicated husk of what it once was.
By bombarding it with bots like Reddit? Honestly, I can’t tell what is bot and not. Are you a bot? Am I a bit? Who knows?
Yeah once the mainstream is fully saturated they’ll turn their gaze to smaller, more “authentic” places. Imagine if they could have a boot infiltrate a tiny little private chat between a dozen people by building up the appearance of an organic presence on the fediverse / indie web.
Now advertisers (and potentially, malicious state actors) have a little spy and influencer in your group.
Gonna get real weird soon when they run out of rich new training data and they begin to consume their own shit. When that happens their entire model will collapse and if the bubble hasn’t popped already that may be what causes it.
Why do you think there has been such an emphasis on hacking with LLMs lately (especially by OpenAI). They figured out all those vulnerability databases were an excellent source for training material. In the past they scraped those, but just for general language training. Now they’ve specifically trained the models on the information within. Some team figured out how to use that data to train a model and have testing scenarios automated so they could write a good reward function. It wasn’t that they figured out the models are good at hacking, they ran out of content and found a new source of good data.
With all the books they’ve been scanning I wonder if the next thing is going to be a writing assistant or editor or something like that. Even though writing good books is an art form and the actual writing down of the words is the easiest part (still not easy tho).
These companies are starving for content and they’ve not just poisoned but absolutely destroyed the content well that is the internet. Given they were already hitting diminishing returns hard, it doesn’t matter too much to them probably. But more compute and storage has also been hitting diminishing returns hard and customers are complaining about the cost. So they are getting a bit desperate on how to improve these things at all.
model collapse is the endgame. that’s the whole point of LLM.
What do you mean? Why would collapse be the endgame?
in a manner of speaking - you always end up there. not by design though. models operate via continuous refinement and you can only optimize a model so much until it is a mess and you need to figure out where to roll back. so you either get shit like semantic drift or variance decay and you can whack a mole it to an extent but then you hit the rlhf wall when the model starts gaming its reinforcement framework and the fat lady sings.
gaming its reinforcement framework
Well put. Never thought of it that way.
the only more or less workable way to keep it under control is maintaining a closed loop small-scale environment - kinda like NotebookLM where you upload documents and that’s all there is to work with - outside of that it is a mess.
Pretty sure they mostly use the pre-AI internet (that they scraped and kept) and synthetic data currently. Probably trying (and failing so far or we’d have heard) to adapt to using video as training material at the moment, but developments there will likely apply to robotics at some point. Here’s hoping the current chuds have crashed and burned before then and that some sanity has taken over from unfettered capitalist oligarchs dreams of computer slavery.
Why don’t you thinj they have not adapted to video as training. Is that not what Flock does?
Flock does pattern recognition, a quite old piece of machine learning. Nothing to do with training a large language model or other ‘AI’ model.
They almost certainly use that to train AI. Big tech is aggregating the data from many places.
Nah, it’s in the name, Large Language Models, what gets marketed as ‘AI’, are trained on text. That’s why they’re ripping up secondhand books (and copyright law, but that never stopped them) at the moment. No one has really cracked Large Vision Model yet, although there are some primitive versions being attempted in robotics, mostly they translate to text, which is why general purpose robots aren’t a thing yet.












