this post was submitted on 16 Sep 2026
70 points (97.3% liked)
Explain Like I'm Five
22330 readers
21 users here now
Simplifying Complexity, One Answer at a Time!
Rules
- Be respectful and inclusive.
- No harassment, hate speech, or trolling.
- Engage in constructive discussions.
- Share relevant content.
- Follow guidelines and moderators' instructions.
- Use appropriate language and tone.
- Report violations.
- Foster a continuous learning environment.
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
It’s not necessarily the big companies themselves doing it? But, there’s a push right now to scrape as much data as possible to feed the training of the AI models, and if your business model is entirely based on selling that data you don’t really care who you’re going to piss off. So, in turn these new scraper bots basically behave in the same way that attack bots do where they’ll be behind VPNs and switch IPs and stuff if you block them.
This is incorrect as you don't need vertical scaling for broad crawling which by definition is a horizontal operation. So you can slurp the web with a per domain throttle that has a respectful limit. The only pressure is browser cache time as if you're crawling a domain you want to take advantage of as much cache as possible but modern websites have very long cache times, like days.
So it's mostly incompetence. Crawling is actually really hard because web is just very complex.
The entire point here is that the bots are being disrespectful.
I'm just correcting you architectural definition. No crawler needs to vertically scale as you define it here and in fact it would be a worse tech for "slurp all data" goal.
The real answer is much more simple - incompetence. The inability to write a good cralwer by leaders in software tech is much damning than anything else. It's like leading car manufacturers not being able to make the wheels round - it's just sad.
However this actually shows a much more important problem with the web - we have no way to defer or distribute negative externalities and people have been trying to solve this since the inception of the internet itself. It'll never be solved but minimized to the point where the impact is just something you tolerate because it's simply impossible to solve and anyone who tells you otherwise is selling you something.
Source: I've been working on similar problems for the last 20 years at least.
You’re not correcting anything, dude.
Whatever, stew in your ignorance then. Bye.
Been stewing at it for longer than you have, buddy.
But WTF is the point of scraping the same data over and over multiple times a second before it even has a chance to change? That's just a waste of resources even on the scrapers' part, because that bandwidth could be used grabbing some other new page instead!
They’re not trying to do that in particular, but on a site like lemmy the same content might be linked a dozen different ways:
https://lemmy.world/comment/25863752 https://lemmy.world/post/51995780/25862545 https://lemmy.world/post/51995780/25860801 https://lemmy.world/post/51995780/25858204 https://lemmy.world/post/51995780/25857153 https://lemmy.world/post/51995780/25857033
There's canonical url tag or 3xx redirect so a decent crawler can resolve this and abort duplicate operations.