If you have copyright juggernauts gatekeeping the data through legislation, then you can't have open-source models.
I'm surprised how many want to shoot consumers in the face just to protect mafia like companies like Universal Studios and YouTube.
This is a most excellent place for technology news and articles.
If you have copyright juggernauts gatekeeping the data through legislation, then you can't have open-source models.
I'm surprised how many want to shoot consumers in the face just to protect mafia like companies like Universal Studios and YouTube.

Remember how china said that a few months ago and the US was all like "communism is stealing our data"
All I'm hearing is piracy is legal now 🏴☠️
The main challenge I see is that Aaron Schwartz and countless others have been prosecuted for access to information but AI companies have been rewarded. There is precedence that access like this is not legal. Had they gone through a library system or used a mechanic like that it may have worked but from what I understand, they used torrents and other mechanics to access the data. So you have companies that go after individual infringement but pursue their own mass infringement. Whether AI generated materials is infringement is above my pay grade but their consumption of the materials seems pretty straight forward as infringement.
Schwartz was saving and distributing copies against the terms of the agreement by which he was able to access journals. What happened to him was heinous but it was pretty dissimilar to how models train on data. And the tormented material was Anthropic, which resulted in the largest copyright settlement in history. Because it was piracy. They briefly tried an argument that their intended use made it fair use, but...that's never how literally any of that worked.
The NY Times and Pearson, two of the most valuable US publishers, each have market caps of about $10 billion.
Let's say Pearson went after OpenAI. They devote an unlimited legal budget to the fight. OpenAI is hoping to IPO as a trillion dollar company. If Pearson went after OpenAI, rather than fight them in court, a deal could be reached first. If that didn't work, if a deal couldn't be reached, OpenAI could bypass the problem completely:
OpenAI would likely have to spend some money on net. The value of Pearson stock would likely be a bit lower after effectively giving away the rights to their works as training data. If signing a durable rights contract the new owners can't escape isn't practical, buying the company and simply holding it indefinitely would also be an option.
But paying and asking for permission first would have been costly at the start and slow. They made a decision, maybe at the beginning, maybe as they realized ethical would mean lost time and placement in the race, and they said screw it, let's go. They also sidelined any AI safety research they were doing (some were making an effort on a difficult problem, but again, $$$ wins).
Project Hail Mary but not for the good of all ...
Project Bloody Mary
Did kegsbreath leak his codename for the next military conflict??
lol
We’re in danger
So hear me out.
From a purely strategic standpoint, the ultimate goal is artificial general intelligence. Eventually, humanity is going to get there. The question is who gets there first, and what they do with it.
The country that achieves genuinely transformative AGI first could potentially gain an enormous strategic advantage. And if AGI is capable of significantly accelerating AI research itself, that advantage could compound very quickly: AGI helps develop better AI, which accelerates scientific and technological development, which produces even more capable AI.
Whoever gets there second may not necessarily be permanently incapable of catching up, but they could find themselves facing a massive technological and economic gap that becomes increasingly difficult to close.
And here's where I think things get uncomfortable.
The countries competing to develop this technology are not necessarily going to have the same restrictions. Some countries may be considerably less concerned about copyright, training data, ethical restrictions, censorship, or what subjects an AI is permitted to discuss.
I'm not saying those concerns aren't legitimate. They absolutely are. But from a purely geopolitical standpoint, there is a potential danger in one country imposing significant self-restraints while its competitors don't.
Because if AGI really does become as transformative as many people believe it could be, this isn't just another technology race. It could become a race over scientific advancement, industrial capacity, military technology, economic productivity, and potentially the ability to develop even more advanced AI.
And, honestly, I would rather the United States and other democratic countries reach that point first than China, North Korea, or another highly authoritarian state.
That doesn't mean I think democratic countries are incapable of abusing powerful technology. Of course they are. The point is that who controls a technology this powerful matters enormously.
So when people talk about AI development purely in terms of corporate profits or whether AI companies should be allowed to train on particular datasets, I think they're sometimes missing the much larger geopolitical picture.
We're potentially talking about a technology that could fundamentally alter the balance of power between nations.
That's a very different conversation.
Disclaimer: I’m not necessarily advocating for AI or LLMs here. I’m simply trying to look at the situation from a purely strategic/geopolitical perspective and consider how governments and other actors might view the race toward AGI.
“Trump‘s cronies suck off the billionaires yet again, to the shock of nobody”
this is why musk got all the CSAM to train grok on, they needed to win that war!
This is the correct decision, just for the wrong reasons.
Rare AI bro W. Fuck copyright.