this post was submitted on 25 Jul 2026
404 points (97.6% liked)

Programming

27863 readers
316 users here now

Welcome to the main community in programming.dev! Feel free to post anything relating to programming here!

Cross posting is strongly encouraged in the instance. If you feel your post or another person's post makes sense in another community cross post into it.

Hope you enjoy the instance!

Rules

Rules

  • Follow the programming.dev instance rules
  • Keep content related to programming in some way
  • If you're posting long videos try to add in some form of tldr for those who don't want to watch videos

Wormhole

Follow the wormhole through a path of communities !webdev@programming.dev



founded 3 years ago
MODERATORS
 

Would be a terrible shame if lots of people opted out.

If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.

(page 2) 50 comments
sorted by: hot top controversial new old
[–] G_M0N3Y_2503@lemmy.zip 13 points 2 days ago (2 children)

Pretty sure all my repos are MIT licensed for the betterment of everyone, but I'm not on the list! So I guess I'm not good enough, or they are failing to follow the attribution clause of it.

[–] Jestzer@lemmy.world 10 points 2 days ago (2 children)

Strange because all of my MIT repos are listed - all the GPL ones are excluded.

[–] stsquad@lemmy.ml 8 points 2 days ago (1 children)

Mine are all GPLv3 or forks of other repos and they are listed. I'm sanguine because the license allows for study and if it's good enough for humans I don't see why it's not for clankers.

[–] fruitcantfly@programming.dev 6 points 2 days ago

For me, and a few other projects I checked, it only has non-GPL repos. But it also does not have everything that isn't GPL, despite the repos being much older than the cut-off date. But it does have repos without a license, which they are simply not allowed to copy.

I wonder if those repos have been deduplicated, and one of the forks (on some other person's account) is included instead. Unfortunately you can only search the first 5M records via the website, and I don't have time to play around with the API at the moment, so I could neither confirm nor deny that possibility

[–] qaz@lemmy.world 2 points 2 days ago* (last edited 2 days ago)

My list includes GPL licensed repo's

load more comments (1 replies)
[–] Kissaki@programming.dev 3 points 2 days ago

From the linked webpage readme:

2.1 Classify. Each file is labelled permissive (at least one permissive license detected, no conflicting non-permissive license), no_license (no licenses detected, or only non-license legal texts such as CLAs), or non_permissive. The permissive allowlist follows the Blue Oak Council list plus licenses categorized as Permissive or Public Domain by ScanCode. Files classified as non_permissive are excluded from both released datasets.

From https://www.bigcode-project.org/docs/about/the-stack/:

v1.1: The three copyleft licenses (MPL/EPL/LGPL) were excluded and the list of permissive licenses extended to 193 licenses in total. The list of programming languages was increased from 30 to 358 languages. Also opt-out request submitted by 15.11.2022 were excluded from this ersion of the dataset. The resulting near-deduplicated dataset is 6TB in size.

So MPL/EPL/LGPL are already not part of the dataset.

So… why were they in there? Was this added for v1.1?

"one permissive license" - So if my project includes a lib and I include the license file for that for the license notice…?

[–] goatbeard@beehaw.org 14 points 3 days ago* (last edited 3 days ago)

Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners

[–] Swuden@lemmy.world 6 points 2 days ago* (last edited 2 days ago) (3 children)

They’re not able to scrape private repos, surely?

load more comments (3 replies)
[–] hexagonwin@lemmy.today 7 points 3 days ago (10 children)

tbh i'm thinking this alone isn't that bad from an archiver/datahoarder perspective

[–] cecilkorik@lemmy.ca 6 points 3 days ago* (last edited 3 days ago)

As long as the datasets are open, it is our best hope. I know it doesn't compensate the people whose work's copyright and licenses have been violated, but I think it's the only realistic hope we've got of getting out of this informational dystopia with a reasonably intact library of humanity's knowledge that hasn't been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.

load more comments (9 replies)
[–] jbrains@sh.itjust.works 6 points 3 days ago

Good thing I never finished a project!

load more comments
view more: ‹ prev next ›