this post was submitted on 01 Oct 2026
544 points (97.7% liked)
Linux
15106 readers
428 users here now
A community for everything relating to the GNU/Linux operating system (except the memes!)
Also, check out:
Original icon base courtesy of lewing@isc.tamu.edu and The GIMP
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
IBM Granite (EDIT: ~~and Apertus~~) do disclose all their training data and claim that all their training data is effectively free of copyright (highly permissively licensed). I have not been able to verify that, due to my lack of skills with the conventions and tools of LLM / Agent training and publishing.
So, yeah, probably (EDIT: ~~two~~ one).
The Apertus Swiss AI unfortunately doesn't seem to live up to its claims, and they have been silent on the issue brought up there. I honestly suspect the same of IBM's Granite, but have not investigated.
Thank you for the link! It does look like Apertus itself might be Free Software (the U.S. copyright office says training can infringe, but is usually fair use), but it can still output derivative works of copyrighted inputs that might prevent them from being distributed as-is (for example, requiring attribution) -- at all, much less under a strong copyleft.
for those who said the deets, thank you!