this post was submitted on 25 Aug 2026
373 points (97.9% liked)
Technology
87572 readers
3603 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Yes, you are acting unethical. Your self-hosted model is neither open nor was it trained with consent. Using them is just as unethical as using the models from OpenAI.
CC BY-SA gives freedom of use. Why do you think additional consent would be necessary? Or do you mean that because the restricting conditions are not met, additional consent would be necessary?
Given the form of transformation, Share-Alike is certainly difficult to meet, if not impossible. Is that what you referred to with "is not open"? Or do you mean something else?
A license may allow something but that doesn't make it ethical to do. At the time these licenses where written the wholesale pillaging of the commons by AI companies wasn't a thing. For that reason I don't think it matters what license anything was published under. Unless you where given permission by the author its not OK to train your model on it.
I consider Open Weights models only open in name, the training data is almost never released. As such it is just an attempt at open washing. Its like releasing some executable files and calling that Open Binaries.
Are you aware of "Implied lack of privacy?" It means places where you cannot reasonably assume privacy, like a public beach or supermarket. Public forums and public wikis are about as implied lack of privacy as it gets. If you post a guide on stack exchange on how to solve a programming issue, you revoke your rights in how that information is used. This includes using it to create new software. If you decide you don't like a program I created using your guides, well... That's too bad for you.
And no, this is not victim blaming, because I'm not defending actual theft or telling people to protect themselves more. I'm stating you cannot steal something that has been freely offered to the entire world.
I don't know why you are talking about privacy. Expectation of privacy is unrelated to copyright. Anyway the point I am making is that a license allowing something doesn't make it moral. Sure, the law allows someone to train their slop generator on my code but I can still think they are an unethical piece of shit for doing so. The same goes for the people who then use such a model. Which is also why I don't feel bad about techbros crying about how they are bullied online for using AI.
Because if you post a public message on a public forum that everybody can see, you don't get to decide how your words are used. That's the fair use license. It doesn't require additional consent.
And how precisely is it unethical to train an LLM on public data? You keep making that assertion, but it's built off of a supposed "Lack of consent." When the fair use license was brought up, you seemed to retreat to the concept that using fair use data to make fair use software freely available to everyone is somehow bad?
You need to explain how the process itself is bad, not just use guilt by association. It seems more like you're building an opinion based on vibes, not actual reason.
Well for once I don't think training a model should be considered fair use. At the time when people released their art and writings under such a license there wasn't the threat of that resulting in them potentially losing their job. I for example stopped contributing to open source completely because of that threat. The reason why I consider training a model on these things immoral is that millions of people put in the effort of creating code or art in various forms and some rich assholes come along scrape it all and sell it back to us while creating ungodly amounts of pollution. It really doesn't matter to me that a few bootlickers run some inferior models on their own system. They are legitimising AI by using it and are encouraging others to also do so. Meanwhile the damage is already done with artists being now accused of publishing AI-generated images because they wanted to do a good thing by releasing their art online for free and the models happily imitating their style.
Okay, then you're upset about copywrite infringement and the resale of the models. I'm not talking about those. You made the universal claim that the technology itself is bad. I'm saying its the PEOPLE who are acting unethically that should get your hate.
You are actively antagonizing people who should be your allies because you've decided carte blanche to hate it all, convincing yourself there's no possible way any of it could be legitimate or ethical.
What about models trained on research papers? Those are LITERALLY designed to spread knowledge.
Are you really going to call me amoral and fracture the opposition to blind greed and destruction, just because I experiment with these local models? What about the scientists that use custom LLMs? Or the research teams?
Some models do use open datasets and largely reproducible training, with the Nvidia Nemotron series being the highest profile example. You are right that many are opaque about training data, though, and this is not okay.
And I vehemently disagree with this.
OpenAI is a different order of magnitude of moral depravity. Thats like saying using Lemmy is just as bad as using Facebook or Palantir.
Lol, tell me you don't understand anything about the technology without telling me you don't understand anything about the technology. Nobody here is defending scraping artists libraries or using copy written material. Go find a different strawman to cry at.
"You just don't understand it" was also what the crypto people always told me when I called their obvious ponzi scheme a ponzi scheme.
Good for you. Keep wielding that shield of ignorance indescriminantly. Do you also think cryptography itself is harmful, or are you focusing on the specific use and harm caused? Guess what, the technology that made the crypto stupidity possible was used for far more than just gambling on imaginary coins. Once again, you're conflating the worst implementation of a technology with its use as a general rule. Would you defend a small set of positive and ethical uses of LLMs? Or do you condemn the entire tech, on the assumption that it is inherently bad? Do you know precisely where and how the tech developed for crypto is helpful to you, and improves your privacy and safety?