I usually don’t even understand the lingo they use. “Open-weighted” is the most recent one, then it usually goes down to specific “models” that everybody is supposed to know about.
These are my thoughts (I will stick to the vague “it” for now, but of course therein lies another question: “and how does all this apply to various specialised AIs”):
- Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
- If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???
edit
Thanks to all who answered.
I guess it’s my fault for asking several questions in one, but this thread has attracted exactly the type of people I’m writing about; several even used the term “open-weighted models” without explaining it.
Asking to get arguments explained, I got more arguments instead.
If you can play modern games, you can probably run a local LLM. The more VRAM the better it’ll go.
For ethics, there’s two issues. The first is power consumption. Running a model locally isn’t consuming any more power than running a game that fully pushes your system. It’s not really a concern.
The second issue is content generation. It’s all created through the theft of content others have made. There’s no getting around that. As far as I’m aware there are no ethical models that specifically have not stolen content. If you’re just using it to chat with or whatever, I think it’s fine. You probably aren’t taking work from anyone, at least to any degree more significant than piracy. If you’re generating stuff to be distributed, then you’re distributing stolen work. If that’s ethical or not is up to you to decide.
For ethics, there’s two issues.
No, there’s seven main issues.
I feel like this question is asked in bad faith if you are able to list seven “main” ethical issues here. I don’t know if I could list seven different issues that aren’t just breaking down power consumption and content theft into smaller pieces, and I’m pretty anti-AI. I guess you could include hardware consumption too. Financing maybe, but that doesn’t play a part in local models generally. Care to explain?
There’s a lot of info that you need to know to explore this space, so I’ll take my own shot at answering. Let me know if anything needs further explaining!
An “open weight” model is an AI model where you can download the data needed to run the model on your own hardware for free. Contrast this with proprietary models-as-a-service like ChatGPT and Claude where you have no access to the data needed to run the model yourself – you can only use it through the services provided, usually for a fee, and which can be taken away from you or changed at any time with no recourse.
The mapping to traditional open source terms does not work well since what you get is a binary artifact.
Those artifacts are released with a license – and many of the models are licensed permissively (e.g. MIT or Apache license terms). You can take those weights, modify them, and then release them as new models – and people do actually do this in practice!
Is it really feasible to run it 100% locally?
Yes. I run models on my own computers and have tried a number of configurations to figure out what works well. The Qwen family of open weight models (from Alibaba) are the ones I’ve found most useful so far. Gemma4 models (from Google) are also useful.
I prefer models that have been modified by the community to remove corporate censorship – i.e. stripping that “As a large language model…” cover-your-ass crap and evasiveness on topics like Tiananmen Square. If that means the model is technically capable of telling me to go kill myself too, so be it; I’ve spent 25+ years dealing with assholes on the internet and can handle abuse from a stupid robot if I have to. (In practice though, they’re usually pretty nice still unless I deliberately tell them to act like an asshole – and then Qwen, at least, starts to sound like a snarky redditor; it’s quite funny most of the time, actually.)
If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
Models do require training to create, yes. It’s generally not clear where the training physically happened IRL – so, yes, some of them probably used power from gas turbines, but others may be drawing power from the Three Gorges Dam in China or solar plants or nuclear plants or whatever else is hooked up to the electric grid where the training happened. Most of them are also not very open about the data sets they were trained on. (There are exceptions to this though!) The Chinese models in particular are almost certainly trained heavily on logs extracted from Western models in addition to using whatever other data they could get ahold of. Whether you think that’s ethical or not is a matter of perspective; how do you feel about Robin Hood?
Once a model has been trained though, it can run on a normal GPU. The power requirements to run an LLM are basically the same as running a video game, or, equivalently, about the same as turning on a few incandescent lightbulbs. (The iGPU in one of my systems uses 100W; the discrete GPU in another system I’ve tried uses 215W under load with appropriate tuning – or 300W if you run it naively.)
If you want to run a model yourself, I recommend using llama.cpp – there are instructions on how to get started with it here: https://llama.app/
These are the models I’ve found most useful:
- Qwen3.8-27B (official version): https://huggingface.co/Qwen/Qwen3.8-27B
- Qwen3.8-27B (uncensored): https://huggingface.co/MuXodious/Qwen3.8-27B-absolute-heresy
- Qwen3.6-35B-A3B (official version): https://huggingface.co/Qwen/Qwen3.6-35B-A3B
- Qwen3.6-35B-A3B (uncensored – ⚠️ mildly NSFW graphics on page): https://huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-heretic
- Gemma4 (official collection – many variants in different sizes): https://huggingface.co/collections/google/gemma-4
If you have an iGPU only, I recommend using one of the so called “Mixture of Experts” (MoE) releases. These are typically named like 35B-A3B or similar; the first number indicates the total number of weights (35 billion) and the second indicates how many are “active” (i.e. actually used during computation) at one time while the model is running (3 billion in the example). These models need less computation to run and stay fast on weaker GPUs. Qwen3.6-35B-A3B is very good in this space and was my go-to model for a long time.
If you have a discrete GPU and enough VRAM, I recommend using a dense model (i.e. one that activates all its weights while answering) like Qwen3.8-27B.
It’s worth noting that people don’t usually use the full quality weights (which are typically ~2 bytes per weight); they use a “quantized” version – compressed in a lossy fashion like a JPEG. Going down to 4-bits (half a byte) on average per weight is about as low as most people like to go – you will see this indicated in names like Q4_K_M. (Quantized to ~4 bits with the K quantizaation scheme, medium variant.) Usually a bigger number is better in the sense of “closer to the original quality” – at the cost of needing more RAM.
Full quality weights are often found as safetensor files on HuggingFace. Quantized weights intended for use with llama.cpp are usually in GGUF file format.
Does that help?
Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
Very useful 35B models need (more or less) 8GB of VRAM and 32GB+ CPU RAM to be usable locally. 16GB RAM might work on a lean system with an RTX Nvidia card and an exl3 quantized model.
I know we are in a RAM apocalypse, but pre-apocalypse, that’s a quite reasonable requirement, IMO.
Personally, I run Deepseek V4 07-31 Flash at 19 tokens/second on a desktop with a single RTX 3090 and 128GB CPU RAM, and that’s an extremely capable model. Again, that’s expensive these days, but pre-ram apocalypse, that is not an unreasonable workstation/homelab.
If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters…
Most open weights LLMs are Chinese. And they are:
-
Trained on pennies. They have to be, as they simply do not have a sea of GPUs like Big Tech. Training costs for their large models are in the millions or tens of millions; a single steel forge has used more energy than all of those training runs combined.
-
China relies more on renewables, and I believe the datacenters aren’t so hastily constructed with gas turbines in the middle of cities.
-
And as of now, they are transitioning away from Nvidia GPUs. Some labs already use Huawei accelerators.
…and stolen IP and stolen personal data?
Yep.
This is a huge caveat.
You can avoid this. Nvidia Nemotron models, for example, are trained on completely open datasets you can download and inspect yourself: https://huggingface.co/nvidia
They are very good, but just behind state of the art.
But in practice, the SOTA models most run use private datasets. Lord knows where the Chinese get it from, but given some common quirks between models, at least some data sources are shared (and possibly government provided?)
…However.
I would argue providing the result of the training as Apache licensed weights counts as “fair use,” in the same way non commercial fan works do.
They aren’t making a dime off releasing those weights. I’m not trying to sell anyone anything when I use them. Where is the IP theft if money isn’t changing hands?
Now, the Chinese LLM services they charge for? I have no excuse for that. Once money is on the table, it is definitely IP theft.
completely open datasets
Not completely – it is mostly open, but they use a dozen or so private datasets for things like training on global regulations, minesweeper (for some reason), etc. To their credit, they do indicate this on the model cards, but it’s not entirely clear what is in those datasets either.
(I fell for that bit of marketing myself awhile back.)
Where is the IP theft if money isn’t changing hands?
Money not changing hands is pretty much the definition of theft.
Unlike the definition of heft which is a photo of OP’s mom.
What I’m saying is it’s akin to writing a fanfic or making fanart of your favorite franchise. Or getting inspired by a painting you see, and making something similar yourself.
Do that for your personal enjoyment? That’s fair use, under the law.
But the moment you start trying to sell it is when you get in legal hot water, and when it’s indeed morally problematic.
The scale is different, but I’d argue a similar principle applies: if you use some model trained on public works from a protected IP, but the model and its outputs are not resold, nor profited from, it’s not theft. The point is beyond money not changing hands; there’s no profit being made from the original author’s stuff. They aren’t being taken advantage of any more than someone viewing their public stuff for free, or someone creating derivatives from private work.
But all that is off the table the moment profit and distribution is involved.
-
Is it really feasible to run it 100% locally?
Yes. My Macbook Air (M2) released in 2022 can run many publicly available LLM models. The ouput is not as fast as using a large powerful datacenter, but for my local needs, I’m not in a hurry. I get about 17 to 30 tokens per second speed running 100% locally.
If yes to the previous: the software doesn’t come from nowhere
For Mac users, the interface comes from here . For the specific LLM models that is a separate question for each.
and ultimately still relies on gas-turbine-powered datacenters
DCs generating power on-site is a relatively new phenomon because existing grids are at capacity, so the only way to bring new DCs online is locally generating power at that DC, usually using gas turbines or even worse, diesel generators. Most if not all of the publicly available models for running on your own hardware were built before those gas-turnbine-generating DCs were a thing.
The public models people are running now have existed for a number of years are likely made on regular utility grid power which is whatever that nation and region uses.
and stolen IP and stolen personal data, no?
The Llama LLM is made by Meta, so probably yes for that one. Deepseek is from an AI research lab in China. QWEN is from Chinese company Alibaba. We don’t know for sure the inputs that created the Chinese models. US AI companies claim a number of the Chinese models are derived from American LLMs, but I haven’t seen (or looked for) proof of these claims.
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically
Likely they mean because you don’t have to pay a large LLM owner in a rent-seeking model to run LLMs, nor does a person’s use contribute to further development by those companies.
or true to FOSS philosophy, because … ???
With the open-weight models it doesn’t rely on a commercial license to use, and the interfaces can be truly open source.
OK but what does open-weight model mean?
“Weights” in LLMs are the “final answer” numbers from the results of model training, and these are the engine of the LLM model. Lets use chocolate chip cookies as an analogy.
The chocolate chip cookies are produced with ingredients, a recipe, labor effort to produce the dough, and then cooking energy/effort to bake chocolate chip cookies. In a traditional AI company, the company gets the ingredients, they write their own recipe, do all the dough creation, and then the energy/labor for baking, and charge you money to get chocolate chip cookies.
An open-weights AI company got all the ingredients, wrote their own recipe, labor effort to produce the dough, and instead of charging you for the dough, you get as much uncooked dough as you want forever. Your only task is to take the dough and cook it yourself and you’ve got free chocolate chip cookies. You can make as many cookies as you want with your own oven. However, you are not given free raw ingredients, nor are your given the recipe to alter it in a way you might like. You can only get the dough for free.
So open-weight AI models (uncooked dough) are LLMs you can use on your own computers (oven) for free and have LLM output ( chocolate chip cookies), but you don’t get the training data (ingredients) nor the training parameters (recipe) that built the open-weight model.
Thanks, that was properly eli5’d!
It means the weights (the numeric data that makes up the model) are publicly available to download and use for free.
In some cases there are conditions such as, if you run the model as your business, you need to purchase a license. That’s usually the largest, most powerful models, though, not those most people would be able to run at home.
Question, what do you do with it and how is the data that underlies the model harvested?
Question, what do you do with it
I use it for both fun throwaway stuff as well as productivity tasks for learning.
Example of fun throwaway:
Do you remember that episode of Seinfeld where the Kramer starts using Facebook marketplace and starts buying the most worthless items before being robbed when trying to get a too-good-to-be-true sale? No? Because it never happened, but you can plug that premise into an crafted LLM prompt and it will pop out a whole TV script with in-character dialog for each actor as well as use of popular existing sets.
Foreign language learning:
I’m studying a foreign language and want to interact with just the level of vocabulary and grammar I have knowledge of right now at my level for practice. I can prompt the LLM to limit itself to just what I know now and adjust the conversation level so I can practice. If I ever get stuck, I can ask the LLM to explain the grammar usage or vocabulary choice.
and how is the data that underlies the model harvested?
I’m not sure what you’re asking here. Are you asking, for example, how the Deepseek model was trained? If so, I answered that above. If not, can you rephrase your question?
You run a program like llama.cpp that can use the weights to run the model, which you can then use for whatever you’d use a model for: figuring out tech stuff, coding, etc. It’s a bit involved, but there are tools that make it easy to get started. LM Studio for instance.
In some cases, the lab that created the model publishes the dataset that it was trained on, and those are usually made of publicly available data. In most cases, though, the labs don’t give details, but the answer likely involves siphoning every web page they could find.
From that perspective local LLMs sound more like classical piracy. Not ethical, not FOSS, but out of the hands of greedy corporates.
they are also doing distillation from the big models from OpenAI & Claude, so open-weight models with similar capability can be available for free and also reduce the big AI companies’ ability to profit from stolen data.
What are the most successful/useful distillations you can run locally?
It’s the backend stuff I wanna figure out so it can convert plain language requests and notes into personal assistant tasks while being flavorfully bitchy about all of it. Apparently RAG and agent tools have something to do with it.
I’ve heard good things about Qwen.
Right… I meant like which specific file.
Some of the open models claim not to use pirated content. I don’t necessarily believe it, but its being claimed.
But also out of the hands of the small artists/devs they’re pirating.
Its much less targeted piracy than the traditional, and its almost weird to see the rules pirates implicitly respected being broken.
You can run them locally, yes. There are models that can even run on phones, but usecase is limited. But it can only be considered ethical, if the training data used is listed or ethically sourced IMO.
AI bros on Lemmy will disagree with me, but most open weight models are still trained unethically i.e, theft. Most proponents of LLMs (who I talked to on bsky), who say local models are ethical, don’t fucking use it. They’re larping on socials about how awesome it is, but none of the ones I talked to are using it in their projects. They mess around, realise it is not as good as the “unethical” options, go right back to Claude
Open weight models Qwen, deepseek, mistral, and the Ollama stuff etc are unethical in normal people’s eyes, but “ethical” enough for AI bros.
From what I searched, there are very few that can be considered ethical - Olmo, Apertus, Starcoder(?). But idk anyone who uses these. My friend at IBM said they used Apertus, but it was nowhere near good as ChatGPT, so they no longer use Apertus now. And these models require minimum 6-8 GB VRAM for their lowest parameter model iirc.
Even the open-weight model bros are lobbying to redefine what ‘open-source AI’ means. That should give you a fair idea about people behind open-weight as well
I would unironically argue that a model primarily trained through distillation of closed frontier models, and then released open-weight with an open-source architecture, becomes “ethical” again.
Something something Robin Hood
Rob the poor’s money from the rich and keep it for your community?
Sounds more like feudal warfare than anything, I can’t see any harm to artists being reduced at all
Just as I suspected…
Thanks for taking the time.
there are very few that can be considered ethical - Olmo, Apertus, Starcoder
This is software meant to be run always and completely locally?
Sorry to whine, but so far nobody has eli5’d what “open-weighted” means, or “model” at that… please?
This is software meant to be run always and completely locally?
That’s what is claimed, I haven’t run them locally since I don’t have a good system.
To be honest, I’m not sure if I can eli5 weights and models, but I’ll try. Think of a model like the base - for example, OpenAI has different models like Astra, Sol, etc. These are different models, like different versions of a software or operating system like macOS, but for AI stuff. Like one would download a software, you download a model to perform tasks.
Weights are vales that can influence inputs of these models to get a desired/better result. What most of these models are doing is mostly predicting what might be the next appropriate text/data to the question you asked. When you ask these AI models what 2+2 is, it is not performing a math operation like a normal program, it is looking at its training data to see what the closest option might be. It is doing pattern matching.
These AI models inside can be thought of like an interconnected network, like neurons in our body, that keep passing information to the next neuron and to the brain to make a decision. (Before understanding LLMs it would help to understand Neural Networks first). These AI networks need weights and biases. These networks perform calculations and weights are used to determine how much importance/weight each input can have on the output. Bias on the other hand, is used to shift/change the output so the AI model can ‘learn’ to pattern match better.
What open-weight models, do is they make the model available for download along with the weights. No information is given on training data. Like with ads, ones with most data emerges victorious i.e, has a better model. So these companies do theft, don’t list their training data afraid of getting caught. I forgot which one, but either Deepseek or Qwen (both open-weight) was caught ‘stealing’ from Claude (not open weight). You can probably guess how much these companies value ethics.
I’m not sure if this entire thing goes away, but local models might be the ones left standing when this bubble pops.
Open weight is different to open source. Open Source AI as it stands, the definition requires a model to have entire thing made public - so the weights, biases, training data used, the model. Apertus, Olmo etc are mostly meeting open source AI definition.
If you need to know more, this is what we’d use to refresh our memory before exams :)
I probably might have made mistakes here, English isn’t my first language either. But I hope you get an idea about these terms
If you really need to understand this tech more, I recommend watching ‘AI for Everyone’ course on Coursera from Andrew Ng. It is free to audit, my friends who took his course were hyped (I wasn’t really interested in AI)
Thanks a lot.
I did not know it was possible to influence AI software/models and nudge them in a certain direction.
These AI companies have even more power than I thought for a long time.
Borderline strawman there but I’ll bite.
Open weight models trained unethically are unethical. Closed weight models trained unethically and then sold back to you for a profit from gas-powered datacenters funded through Ponzi schemes are substantially more unethical.
From there it’s a harm reduction calculus. No, those are never pleasant.
So do you let the closed weight labs conquer the field unopposed just so you can feel better about yourself? That’s a valid stance, FWIW, and it’s also super easy and convenient because you don’t have to do anything. It especially makes sense if you believe it’s still possible that AI will just go away on its own. I don’t, myself, not anymore, so I encourage the use of open weight models, however grudging, so people don’t give money to the closed labs and in the worst case aren’t eventually stuck with the maximally unethical options. And we’ve not even touched on the nightmare labor replacement scenarios that seem every day less unlikely. Am I right? I have no clue. Like you, I’m just trying to make the best choices I can in a world that’s gone to shit. I’d recommend dropping holier-than-thou attitude either way, though, because it doesn’t help our side. Man.
They are all trained on copyrighted material without permission, no LLM is ethical
What’s not ethical is a copyright system thag enforces artificial scarcity where there is no need for it.
Piracy is not stealing, and is not inherently unethical.
Everything is copyrighted regardless of it being owned by a corporation or blogger
I don’t partially care about the former
The absolutely most generous copyright protection length is in Mexico which is 100 years plus the life of the author when created. So lets generously say a total of 160 years. Most of the rest of the world is 120 year max. Anything outside of that is in the Public Domain. Many things had a much shorter path to the Public Domain falling into it in as little as 20 years.
So lots and lots of stuff is not copyrighted.
You can also train your own local models with license free material if you wish! I think one of the easiest ways to get into that is by using software like unsloth (that’s the one i am using), an open source no-code tool which can be both used to train models on whatever data you wish and to run models either locally or using an inference provider.
Quick example for something like that which is also not dependent on copyrighted material is RAG, where you can provide the 400-page manual for something and then can chat with “the document” to get explanations, ask quick questions without searching for possibly multiple occurrences of a specific term and similar stuff.
Doesn’t RAG require a pretrained model still? Presumably on copyright material?
It does not per se need a model using copyrighted stuff. You can get by using a model without that - basic language skills and technical jargon can easily be trained with open material.
In theory that makes sense, but does this actually exist?
It does if you want to go that route. https://www.kaggle.com/datasets is a good starting point.
Thanks for the info!
Some don’t speak any “language” they are trained to “speak” and “think” in terms of election orbitals and bonding energy. They are used in pharma and materials science to work on intractable problems like superconductivity and meds for Parkinson’s.
that’s not really an LLM then, is it?
But it is AI.
true, but that’s hardly relevant as a reply to s comment which is clearly referring to LLMs and not machine learning in general
It is, cause it uses the same architecture, https://www.geeksforgeeks.org/artificial-intelligence/exploring-the-technical-architecture-behind-large-language-models/
an example: https://github.com/bowang-lab/scGPT “talks” in RNA and is incredibly accurate and able to do predictions
There is Apertus which at least claims to only use permissively licensed sources and respect robots.txt opt-outs. They outline their methodology and source datasets in this document. Though I haven’t checked the actual sources myself.
do you consider “piracy” like zlibrary and annas archive unethical?
If a business is using them, yes.
i agree.
in that case, wouldn’t it be similar ethically if people use these models on their own machine for personal use?
As an exercise so I could learn the technology, I trained a model exclusively on the Public Domain works of L Frank Baum, specifically the “Wonderful Wizard of Oz” series (did you know there are 14 books in the series just by Baum?!). The model created is not useful as a tool, and more often than not just produces English gibberish, but every now and then it can produce original coherent statements.
Again, I didn’t do this to produce a useful tool, but rather an exercise to learn how to build models.
They are all trained on copyrighted material without permission, no LLM is ethical
However, this means I can refute your statement because 100% of the input data is public domain novels. Also, I trained it on my own hardware powered 100% by solar power.
Cool that you did that. However, “Large” as in “Large Language Model” generally refers to the dataset size, and 14 books is not Large.
Human brains are all trained on copyrighted material without permission, no human brain is ethical
You can run model locally, a gaming PC is enough, may-be not for cuting edge models but if you want to generate clues for a RPG or rephrase a letter it’s good enough. You can look for LM studio and Stability matrix for example
Bonus if you’re using solar power to generate your responses.
It is absolutely possible to run 100% locally, but in practice, at the high end, only quantized models. The full size top end models require hundreds of GB of video ram, and while you can buy that, it’s stupidly expensive. Quantized models can often perform nearly as well with a small fraction of the ram, but they do sacrifice a little in precision.
These models (well the good ones) ultimately all trace their origins to what you’d likely consider “stolen” data. Whether that’s ethical is debatable. If you’re in the “information should be free” camp, there may not be an issue here.
As for the power/environmental impact, for what they do LLMs are actually very low impact per-request. If you’re concerned about your personal AI power use, then I hope you never fly in an airplane, and minimize your driving because those are much bigger issues.
It’s the scale of use that makes AI an environmental problem, and that’s a question about corporate use of AI, not personal use of AI.
As for the power/environmental impact, for what they do LLMs are actually very low impact per-request.
Worth noting that a request is often dozens of requests now that there’s “reasoning,” even for search. As I understand it, a model will take a question , figure out the context (one request), reform the question so it yields better results (another request), if it’s doing a web search there’s requests for each result, another to compare, another to check if it answers the original request, if not it loops and does it all over again. So one request is easily, and often, dozens of requests. This is one way Ai companies can say to investors, “see, look at how much usage has increased.”
Also, we need to factor in the power to scrape and train each of those models, build the datacenters which is near impossible as these companies are not transparent about it and actively try to obstruct investigations into it. Then there’s the redundancy of all these different companies competing and doing roughly the same thing at the same time, as fast as they can, so it’s orders of magnitude inefficient energy consuming before it gets its first user prompt.
Comparing it to other assaults on the environment is not only hard to do, but a case of “the worse negates the bad” fallacy.
I do see what you’re saying. You can account for all of these factors and it still turns out that, largely, individual LLM use just doesn’t use that much power compared to most things people do day to day. Inference is so cheap that even dozens of requests don’t amount to much. I could look up and give you a bunch of numbers, but I don’t think that’s likely to convince anyone who doesn’t do the research themselves. It’s so easy to come up with sources that say what you want. I’d encourage you to actually look into this yourself.
Training costs are higher, but you train once and use repeatedly. Right now, total training costs are stupidly high, but that’s because we’ve got an arms race between the frontier labs to spend as much money and compute as they can for truly marginal gains in quality. The solution to that problem isn’t for individuals to stop using AI, it’s to stop those assholes from wasting so much power.
Individual LLM use is so cheap, that it really isn’t worth wasting people’s energies thinking about limiting that. Instead of being distracted by attempts to make this an issue of personal responsibility, we should be focusing on what will actually make a difference. We should be focused on supporting policies that lead to systemic change. A carbon tax would change corporate behavior right quick, and not just for AI companies.
Thanks. What are quantized models?
I’m going to simplify a little here, so don’t take this completely at face value.
Models are, quite literally, long series of numbers (called weights). A model might store the weights in 16 bit numbers (that is, 16 binary digits). The size of the model (and how much memory it needs) is determined by how many weights there are, and how many bits each weight takes. You can take a 16 bit model and rework it to use 8 bit, or even 4 bit numbers. The result intuitively behaves a lot like the same model, but with less precision to the weights. That makes the model take way less space in ram, but also makes it more likely for concepts (encoded in the weights) to overlap, which impacts model quality. Often the effect is that fine distinctions get lost.
Kind of like “compressed”. It takes longer/more effort to run them, to produce the same result as a the same model that has not been quantized, where that non-quantized version would consume significantly more RAM but produce the result faster. You would typically only run the quantized model when you’re starved for RAM, which most of us are running LLMs locally.
Think like zipping a file with file compression. It takes less space, but has to be unzipped for you to have usable files again.
The power impact was something that in the early days worried me. Looking into it I agree with you to some degree. Like using it instead of a search engine I think is by and large a wash. One prompt will likely take more energy but will give you information that likely would have required searching several times modifying the words and jumping between sites which are rendering all sorts of things. Heck If I booted into a command line and connected to an llm Im almost sure it would be significantly less energy. If you chat for entertainment instead of streaming vidoe also lower energy use. Now I think one thing is in making things. It lets people who otherwise couldn’t make pictures and videos and code. In the large majority of cases what is made is going to be disposed even for folks that eventually make something they care to keep around or use. While using software to do the same uses a lot of energy the only people doing it generally where making long lasting things for projects or such. So that is where I question it. Still I will have it make a picture to use in an rpg or such.
I’ve come round to the idea that I will not use LLMs regardless of how “ethically” they can be sourced because their outputs are harmful in a way that builds up, like a very small dose of poison.
I’m not holding this as a hard stance, just a current analysis of the technology and how it might affect my life in my circumstances. I think we’re well into a world where we need to be deciding if some technologies don’t suit our lives, because there certainly are even more harmful technologies to come.
It might seem like a romantic or artist’s approach, but with so much spiralling out of our control I’m trying to make the active choice to be more human.
In leaving progress to the machines, in letting technology go forward on its own terms and selecting from it, with what seems to us excessive caution, modesty, or restraint, the limited though completely adequate implements of their cultures, is it possible that in thus opting not to move “forward” or not only “forward,” these people did in fact succeed in living in human history, with energy, liberty, and grace?
Always Coming Home - Stone Telling Part 3 by Ursula K. Le Guin
I’ve come round to the idea that I will not use LLMs regardless of how “ethically” they can be sourced because their outputs are harmful in a way that builds up, like a very small dose of poison.
How is LLM output any more harmful than output from a human you don’t know? I would agree with you if one were to simply blindly accept anything an LLM gives you as factual. However, I am skeptical of what humans say too. Some are inaccurate out of carelessness, some out of malice. Critical thinking is the key to protection from both errand LLM output and fallible humans.
Additionally, this thread is about locally running LLMs. One of the most dangerous aspects of most large LLMs is the sycophancy where the LLM will try to tell you what you want to hear, even if it needs to creatively invent things that don’t exist or are not true. If you are not aware, running locally means you have all the controls on the model. You’re not subject to whatever settings a large hyperscaler sets up for you. This means, among other things, you can turn down the “temperature” setting, which is the lever that controls how creative an LLM is. Practically what this means is, if you set it to “0” you are allowing it no creative action. If you ask it a question it doesn’t have actual training on, it will tell you effectively “I don’t know” instead of making something up just to have an answer.
It might seem like a romantic or artist’s approach, but with so much spiralling out of our control I’m trying to make the active choice to be more human.
The older I get the more disappointed I get in humanities large group decisions and actions.
I’ll take a crack at it.
I run AI on an old Nvidia p40 purchased online for less than $200. The models I run on it came predominantly from Chinese companies and groups. On balance, they:
-
use much less fossil fuel in producing these LLMs than their Western counterparts
-
produce models that run much better on lower end consumer hardware
As for training data, the inputs used to create these models, it varies greatly but the most popular line, Qwen, Came from the company’s own data from operating such huge networks and systems for so long
I really fail to see how any of that is worse than playing a video game.
None of this takes away from the very real issues around data center build out and Western companies using the systems to scare people and continue an economic bubble. That’s all true and bad. But there’s nuance. Not all AI is created equal
use much less fossil fuel in producing these LLMs than their Western counterparts
Utterly false, given that all the decent Chinese models (especially Qwen) are distilled from Western frontier models.
They literally could not exist without the enormously carbon emissive western models existing first to train them.
Also China has an insanely diversified grid, that utilises almost a global scale of fossil fuels.
So its like green communism at best
distilled from Western frontier models.
What does this mean please (remember the title of the post!) and how do you know?
Basically, it’s really really expensive (from a compute standpoint, and more compute = money and energy) to train a decent LLM just using books and conversations etc. Anthropic, OpenAI, etc spend a truly insane amount of money doing this is in data centers.
Running an LLM like that is also extremely expensive from a compute standpoint, and right now the true frontier models can basically only be run in datacenters.
However, once you have a really good LLM, you can use it to train another LLM. Because it’s basically already refined all its original training data (and is capable of further refining it’d output on the fly), training a new LLM from it is vastly easier and cheaper.
Additionally, you can use a much smaller and more efficient model, so it can run on less powerful hardware.
So all local LLMs that are any good that exist right now, were trained using the outputs of existing super expensive frontier models. They could not exist without the frontier models to train them. This is also how OpenAI and Anthropic create their cheaper models. Mythos/Fable is anthropics current frontier model and they distill it into (use it to train) their cheaper models like Sonnet and Opus.
Wow. as if the original models weren’t rickety enough.
They’re honestly not that rickety. If you’re getting your news from social media (including Lemmy) you may want to actually try them rather than just take the random incidents where things go wrong as broadly representative.
yeah the comment was bad enough I looked at the user. been around for 2 years and no posts or comments till this one. made a note on the account.
-
I have a friend who runs his AI locally because he doesn’t want to pay for tokens.
I don’t know if 5 year olds are allowed to watch 1.5h videos, but if so: this one has all the pro and con arguments explained nicely and makes a point why it’s not a good idea to shame people for using LLMs.
Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still.
There are small models you can run 100% locally on a mediocre computer or even a phone, and I mean you could be 100% offline and it will still work.
The better local models require a higher end gaming GPU or an ARM Mac with a decent amount of RAM, but nothing too extreme. You could get a computer to run them for about $2000 to $3000.
the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
The small local models typically start with a full infustrial-size model and then they “distill” it to smaller models. The concerns over training data are still valid. I’m not sure how much the concerns over power usage are about training vs running the full industrial size models commercially. Running a small local model uses a tiny fraction of the power it takes to run an industrial size model, and the training is a one-time cost (except these companies never stop training because they want to make next year’s model better and faster).
I do think it’s worth pointing out that running a model locally keeps your conversations with it private. Remember, it can be done 100% offline. So your data won’t be stolen.
Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still.
It’s absolutely possible, and done more and more. Smaller models exist, and run on sub 8GB VRAM without any issue. If you have a gaming setup, you can run the bigger models without any issue locally. You don’t need Astra or Mythos or whatever. As a daily driver for “light” tasks (e.g. summarizing a document) small models perform without a significant difference to frontier cloud models. For bigger tasks (e.g. coding or agentic workflows) you need a bit of beef on your graphics card, but it still runs on regular consumer hardware. Anything above that is not needed for any normal use-cases.
If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
This has multiple layers:
- for running models locally, FOSS solutions exist and are no different to any other piece of software
- models themselves are usually trained on beefy data centers and questionable data, but alternatives exist, both for training and curated data sources. Second but - those usually underperform, because LLMs need those massive data sets. Limiting the dataset limits the output quality drastically.
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???
Also multiple answers:
- …because first there is a massive amount of hardliners both in the pro- and anti-LLM camp, and differentiated and objective views are rare. Ethical software and “true FOSS” are also very heated topics currently.
- …because, while possible, objectively non-ethical use of LLMs does happen more (public vibecoding of products, weaponized use in cyber and conventional warfare, use in surveillance, etc.).
- …because there are absurd amounts of money circulating in the AI bubble, it’s too big to fail, and ethical and for-profit usually do not match.










