this post was submitted on 02 Sep 2026
327 points (98.8% liked)

Technology

87773 readers
4063 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] assembly@lemmy.world 8 points 14 hours ago (3 children)

The main challenge I see is that Aaron Schwartz and countless others have been prosecuted for access to information but AI companies have been rewarded. There is precedence that access like this is not legal. Had they gone through a library system or used a mechanic like that it may have worked but from what I understand, they used torrents and other mechanics to access the data. So you have companies that go after individual infringement but pursue their own mass infringement. Whether AI generated materials is infringement is above my pay grade but their consumption of the materials seems pretty straight forward as infringement.

[–] FatCrab@slrpnk.net 1 points 9 hours ago

Schwartz was saving and distributing copies against the terms of the agreement by which he was able to access journals. What happened to him was heinous but it was pretty dissimilar to how models train on data. And the tormented material was Anthropic, which resulted in the largest copyright settlement in history. Because it was piracy. They briefly tried an argument that their intended use made it fair use, but...that's never how literally any of that worked.

The NY Times and Pearson, two of the most valuable US publishers, each have market caps of about $10 billion.

Let's say Pearson went after OpenAI. They devote an unlimited legal budget to the fight. OpenAI is hoping to IPO as a trillion dollar company. If Pearson went after OpenAI, rather than fight them in court, a deal could be reached first. If that didn't work, if a deal couldn't be reached, OpenAI could bypass the problem completely:

  1. Spend $5 billion to buy a controlling share of Pearson.
  2. Fire the entire leadership team and install OpenAI minions in their place.
  3. Once they control Pearson, sign a long-term licensing deal with OpenAI with very generous licensing terms and huge early cancellation fees.
  4. Sell the shares back on the market at a (likely slightly reduced) value.

OpenAI would likely have to spend some money on net. The value of Pearson stock would likely be a bit lower after effectively giving away the rights to their works as training data. If signing a durable rights contract the new owners can't escape isn't practical, buying the company and simply holding it indefinitely would also be an option.

[–] Rhaedas@fedia.io 3 points 13 hours ago

But paying and asking for permission first would have been costly at the start and slow. They made a decision, maybe at the beginning, maybe as they realized ethical would mean lost time and placement in the race, and they said screw it, let's go. They also sidelined any AI safety research they were doing (some were making an effort on a difficult problem, but again, $$$ wins).