this post was submitted on 01 Oct 2026
544 points (97.7% liked)
Linux
15106 readers
428 users here now
A community for everything relating to the GNU/Linux operating system (except the memes!)
Also, check out:
Original icon base courtesy of lewing@isc.tamu.edu and The GIMP
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
i think the copyright question is still unanswered: it could turn out that any LLM-generated code is a copyright violation by definition unless trained exclusively on a clean, legitimately-obtained dataset (which few of the major models are).
Will projects that allow LLM contributions have to roll back years of progress when the other shoe finally drops? Seems like a huge risk, especially for FOSS and copyleft. I think disallowing LLM-written contributions until this is all sorted out in the courts is the pragmatic move from a legal perspective.
If it turns out that it gets ruled as copyright violations, you can bet your ass they're going to reform copyright law instead of rolling everything back.
I don't disagree that it is a risk, and I am trying to move toward projects that do have err on the side of avoiding that risk. I have NetBSD on my laptop, and when I get a little more comfortable with it, I intend to convert the other Linux installations I maintain.
BUT, I believe the BSDs already went through a situation where some of their source was possibly under restrictive copyright and rather than "rolling back", they "simply" identified the possibly infringing code and re-wrote those sections to have the same function (which can't be copyrighted) without sharing any creative expression (which is). So, even the projects that are taking the risk that an LLM (or other generative AI) generates infringing code might not have quite as much cleanup / lost effort as you describe.
Also, LLM out isn't automatically a derivative work of the training data. I'd have to dig through some other messages to find an exact quote from their documents, but I believe they (EDIT: the U.S. copyright office) said only output that is "significantly similar" to training data is potentially infringing. That does further limit the risk.
I still think it's too high of a risk because well-meaning contributors might incorrectly introduce infringing code, since for models that don't disclose their training data (Claude, Copilot, Gemini, etc.) even dedicated contributors don't have the information they need to discover the output is infringing. In that past, that result (introducing infringing code) was generally limited to the acts of malicious actors that are submitting code they know to be infringing to poison a project and open it to legal action.
But, I can't ask that someone (i.e. a project maintainer) substitute my risk/reward judgement for theirs, and I have no experience maintaining a large project. All of my code contributions are to either projects others maintain, or my own hobby projects that I doubt have any users other than myself (and I don't even use all the published/available ones anymore).
The issue isn't necessarily that the work could be considered derivative, but rather that the notion of copyright exists only for work authored by a human. If it's not produced by a human, then the copyright belongs to no one, and no one can license it because no one owns it.
I think this argument is BS, there are several remix/sample based albums that count as derived works and AFAIK, no one is getting paid.
Note that the remix/sample example hasn't always worked out as you state: https://en.wikipedia.org/wiki/Bitter_Sweet_Symphony#Credits_dispute https://en.wikipedia.org/wiki/My_Sweet_Lord#Copyright_infringement_suit
In many cases, "AFAIK" in your case you may have no idea that in fact, the copyright holder is being paid. Or the copyright holder is one and the same, with rights sometimes assigned to someone other than the musicians involved.
https://en.wikipedia.org/wiki/Fair_use#3._Amount_and_substantiality
you wanna make the argument that 5000000 hello world projects are a substantial part of an AI?
Not exactly the same, and the music industry has had plenty of lawsuits going both ways on that kind of thing establishing a status quo for remixes and samples in music
"Most" music is also under a compulsory licensing system, while virtually no code, prose, or visual art is.
SilvaGunner 🫡
It's answered, unless all big tech is going down suddenly, they are allowed, copyright violations are for the poor anyway.
Note that it is mostly 'answered', but only really for the services and models themselves. Part of the defense in the case where Getty showed that it produced knockoffs of the Getty images was that "well, as the operator you are responsible for the output, not us".
So the operator is still on the hook if someone comes along and claims, even unwitting, infringement. In practice for code that's likely a tall order, as the sloperator is likely to keep their so close as to be a knock off pretty private, or the open source developer lacks the resources to realisticly track down offenders.
Closed source companies are more likely to come after folks, but that's a lower risk because they are so maniacly defensive about their code it probably never trained a model.
Few? Do you have some example, maybe two?
IBM Granite (EDIT: ~~and Apertus~~) do disclose all their training data and claim that all their training data is effectively free of copyright (highly permissively licensed). I have not been able to verify that, due to my lack of skills with the conventions and tools of LLM / Agent training and publishing.
So, yeah, probably (EDIT: ~~two~~ one).
The Apertus Swiss AI unfortunately doesn't seem to live up to its claims, and they have been silent on the issue brought up there. I honestly suspect the same of IBM's Granite, but have not investigated.
Thank you for the link! It does look like Apertus itself might be Free Software (the U.S. copyright office says training can infringe, but is usually fair use), but it can still output derivative works of copyrighted inputs that might prevent them from being distributed as-is (for example, requiring attribution) -- at all, much less under a strong copyleft.
for those who said the deets, thank you!
Few could be zero lol, i have no idea.
IBM Granite (EDIT: ~~and Apertus~~) probably qualify.