At first I thought this was going to be about how in IT knowledge kinda has a half life of usefulness.
Showerthoughts
A "Showerthought" is a simple term used to describe the thoughts that pop into your head while you're doing everyday things like taking a shower, driving, or just daydreaming. The most popular seem to be lighthearted clever little truths, hidden in daily life.
Here are some examples to inspire your own showerthoughts:
- Both “200” and “160” are 2 minutes in microwave math
- When you’re a kid, you don’t realize you’re also watching your mom and dad grow up.
- More dreams have been destroyed by alarm clocks than anything else
Rules
- All posts must be showerthoughts
- The entire showerthought must be in the title
- No politics
- If your topic is in a grey area, please phrase it to emphasize the fascinating aspects, not the dramatic aspects. You can do this by avoiding overly politicized terms such as "capitalism" and "communism". If you must make comparisons, you can say something is different without saying something is better/worse.
- A good place for politics is c/politicaldiscussion
- Posts must be original/unique
- Adhere to Lemmy's Code of Conduct and the TOS
If you made it this far, showerthoughts is accepting new mods. This community is generally tame so its not a lot of work, but having a few more mods would help reports get addressed a little sooner.
Whats it like to be a mod? Reports just show up as messages in your Lemmy inbox, and if a different mod has already addressed the report, the message goes away and you never worry about it.
That would make an excellent new shower thought! I suppose some of that would be due to obsolescence, and some from the physical degradation of digital media. We seriously need to get back into the stone tablet game or there will be nothing left to mark our civilization in a millennium than a whole bunch of porcelain toilets scattered around in the rubble.
Apologies if I'm reading too much into what you've said here, but it reads to me as if you're implying search engines couldn't keep up with the volume of the internet and AI is the logical next progression.
Search engines actually worked pretty amazingly at one point, but then were kneecapped because Google et al were more interested in advertising money than useful search results. Then when they needed an excuse to justify ludicrous spending on AI, they shoehorned AI into web search where it is very much the wrong tool for the job if you want accurate, high relevance results.
For AI, it's as though the library were just a massive pile of books tossed around haphazardly and the librarian had to make sense of it enough to pull whatever you're looking for out of the chaos. Is it at all surprising, then, that it takes a huge amount of energy and computing resources to make any of this work?
It's more like the librarian has taken it upon themself (for some reason) to write a book that is the mathematical average of every other book in the library. Then when any patron comes looking for a specific book, the librarian just hands them this slag instead hoping they won't notice the difference.
Now that said, I 100% agree that running a good "old-style" search engine does still take a huge amount of computing resources, but I'd be surprised if the requirements were anywhere near what is being thrown at training LLMs.
It's more like the librarian has taken it upon themself (for some reason) to write a book that is the mathematical average of every other book in the library. Then when any patron comes looking for a specific book, the librarian just hands them this slag instead hoping they won't notice the difference.
Honestly, best description of LLMs for search I’ve seen so far.
Ah, so you're saying it's more of an enshitification thing that killed the search engines? I can believe that to an extent.
But as the amount of content on the Internet has increased, it's also become more difficult to narrow the queries to get what you're really looking for. The main thing for me with AI is it allows you to refine the search without discarding the context of what you have already tried.
If the content were more organized to begin with, though, I suspect you could drill a bit deeper into the knowledge tree and then apply your conventional (non-AI) search over a more relevant branch?
Let's not forget Search Engine Optimization (SEO) and content farms which were actively corrupting (adding entropy to) search engine results even before LLMs, yay capitalism. They've dialed it to twelve with LLMs.
Heh, I had actually forgotten about that. It seems almost quaint compared to where we're at now.
It's not just increase in content. Search engines don't even do searches based on keywords or logic operators anymore. They run some NLP on the query input and do who knows what to give you results. If we could still make searcher that contain specifically certain terms or exclude others, it would be still possible to navigate into the mess
Google has a financial incentive for you to spend more time on their site. If you are gone in one click, you see less ads. Same problem with dating platforms. The user's interest is in conflict with the platform's.
There are leaks/interviews how their search and ads teams were at a conflict internally, ofc money won.
Does the ad revenue account for people who hardly realize the ads are even there?
If you hardly realize they're even there, you're more likely to click on them since most advertisers pay to be shown on relevant searches, not random ones.
E.g if you run a pet food store, you'll let Google show your store whenever pet related things are searched for. And that's where having a blog comes in handy, you can have it show articles relevant to the search rather than just the store front page.
Search algorithms are meant to work regardless of the amount of content or sorting. Page rank used to be a game changer back then to sort websites by relevance.
AI has no search algorithms. It's about 1000% less efficient, because its algorithms are built for token prediction, not sifting through websites.
So yeah, my money is on enshitification.
AI has no search algorithms. It’s about 1000% less efficient, because its algorithms are built for token prediction, not sifting through websites.
It doesn't need to; It just relies on the normal web search to get better data than what it could generate from its training data alone. This wasn't a thing with the early popular LLM interfaces, which is why they were ridiculously bad. They still make plenty of mistakes, but at least now you get a link to what it looked up (so you can facepalm when you see it's way too generic and your question was very specific, so the generated answer is still wrong)
I often get dead links or those that state something completely different from what the LLM says. If we'd just cut out AI as the middleman, oh well. 🥲
Then when they needed an excuse to justify ludicrous spending on AI, they shoehorned AI into web search where it is very much the wrong tool for the job if you want accurate, high relevance results.
AI is very much the right tool for sorting through the high volume of garbage search results looking for nuggets of relevance. Google built it into search because if they didn't, they would be cut out of search other than as an MCP provider. I'm sure they'd rather not spend compute running an AI summary of every single query, but they would be out of business the moment someone else provided that service.
Now is it perfect? Of course not. Verify from sources, but you get to those relevant sources faster and easier with AI doing the sifting because Google search has grown so fucking worthless.
AI isn't better than Google was, but it's better than Google is.
Google is not at fault for search becoming so terrible. Everyone who engaged in SEO is. The internet would have sloppified much sooner if Google and other search engines had not relentlessly pushed back against SEO and continuously modified their algorithms. Lots of people definitely underestimate the extreme negative impact of SEO, and the lengths that search engine operators have to go to to have their service be at least somewhat useful. As soon as you have a search algorithm, people will attempt to manipulate it.
search engines are failing us
But this isn’t a problem with the internet content or search engine technology. This is because every search engine company just wants to show you ads and paid content rather than actually search for things anymore. Which 20 years ago, it used to do quite well.
It's gotta be partly to do with content though, right? Are there any search engines that effectively sort through the slop?
I meant to go back and edit or reply to another comment to walk back my statement a bit. There is a problem with content. The walled gardens hide content from search engines. So you can only find their content by being logged in to their service and using their search tools (which are also shit.) So yeah, there’s that too.
Many people have noticed a decline going back to 2020 when Google decided they need to show more ads, which means you need to search more times or check more pages than the first one rather than getting the desired result right away.
Kagi seems to be better at this since it's paid and doesn't need to show you ads.
But of course you're right that content itself has sloppified too.
yep
Search engines were, at one point, pretty fuckin good at finding what you needed, and would put it at the top of the first page of results
but then some companies coughgooglecough realized that users would be okay going to page 2 or 3 for their results, so they intentionally undid a decade+ of search engine improvements, and started forcefully putting relevant shit 2-3-4 pages in, so they could harvest more clicks and shove more ads in your face.
and thus enshitification began.
you're completely right. there's tons of problems with the way that we organize information today. another big issue is that a lot of publicly funded scientific research is behind paywalls of journals, and the journals don't even have a good indexing system where you could easily search for stuff! it's infuriating.
It seems like this problem stands to get worse too, as researchers see a need to protect their work before an LLM comes along and makes the big discovery on their backs.
Wow, there is a lot to think about in this comment thread! It's going to take a lot more showers.
I'm not an information theorist, but I do wonder sometimes if nature prefers a more chaotic data organization? Even when you look at lower levels than the Internet. Hash tables seem to be winning over tree structures. Randomized network protocols are winning over orderly ones. And even if you look at AI, this used to be a broader term covering other topics such as expert systems which tried to condense human knowledge and a very low entropy way.
If there's an inherent bias towards disorganization, where does it stem from? I guess there can be a somewhat better average performance, even if the worst case can be really, really bad. There may also be better resilience? Organized systems tend to fall apart very quickly when something goes wrong.
The universe/nature does always tend to higher entropy, yes - that's the 2nd law of thermodynamics. "In an isolated system, entropy can only increase".
I posit that the "purpose" of life is to fight back against this inevitable increase - every alive thing is a self-sustaining island of low entropy after all - by processing and creating more low entropy states - "information" in a broad sense.
You are also right in the resilience aspect - by definition, a low-entropy system has a vast number of ways to increase its entropy and "decay", while self-improvement is so statistically unlikely to be virtually impossible.
You might find this an interesting read.
https://www.scientificamerican.com/article/a-new-physics-theory-of-life/
It posits that life may have got its start as an entropy-increasing mechanism. But I still see your point that in its present form, it seems to produce a lot of complex organization.
One missing piece is that librarians are acting as curators of the data. Not anyone can come to the library, put books around, and game/pay the book finding process to make sure their books come first. Having human curators behind an internet search engine is not realistic.
Unless the library gets bought by some oil company and they force the librarian to add a lot of pro-oil books.
Those libraries exist and that's why we publicly fund libraries.
Not anyone can come to the library, put books around, and game/pay the book finding process to make sure their books come first.
Well, as I recall, things weren't quite idyllic in the conventional library scene either. People were always moving books around, making librarians furious. Every field trip to the library always began with that lecture. "For the love of all that is holy, do not put books back on the shelf. Let us librarians take care of that." And once I got to universities with their reference collections, I found that people were purposefully hiding texts in other sections so that they would have exclusive access when they returned. Fun stuff.
I think you're writing stories for librarians. Their whole existence relies on people using the books. They don't want you to put the book away because, A- you'll do it wrong and B- they count them for their general statistics which helps with their funding. They don't want you being a pretend librarian, they went to school for it and are better at it than you.
People hide books, forget to return books and damage and destroy books. Librarians moniter and curate the collection. Libraries are not museums, the catalog ebs and flows and is a constant state of flux. I can tell you that me wife is happiest when she comes home from working a busy shift where the library was full of people.
Sure, there are the librarians that yell at kids and give you the stink eye for asking questions, but they're just everyday-sadists and they're everywhere.
Yes, the internet is decentralised in every way, not only its infrastructure. Even unorganised. If that's your big discovery, I respect that. But...
the library was the main source of human knowledge, and it was very organized through the Dewey Decimal System
Not all of humanity uses the DD system.
Today, even the search engines are failing us, and we are turning to AI.
This simply isn't true, and everything you write after that. It's an insult to actual librarians.
There is no "Artifical Intelligence".
Indexing the internet is a big job. We had data centers even before A-not-I. And people were complaining about how much energy they use. Had we known then how much worse this is going to get very soon....
Artificial Unintelligence does not do that job better btw. It doesn't do the indexing at all. It just wastes energy. And, as another commenter pointed out, it actively makes the situation you describe worse.
Yes, the internet is decentralised in every way, not only its infrastructure. Even unorganised.
What's intriguing to me about the fediverse is it's trying to layer some order over a decentralized physical infrastructure. I think that's so cool and really want this sort of idea to work!
Not all of humanity uses the DD system.
Yeah that's fair. I think even my university was using a different system, and of course there's a lot more out there than book knowledge.
Indexing the internet is a big job. We had data centers even before A-not-I. And people were complaining about how much energy they use. Had we known then how much worse this is going to get very soon….
Artificial Unintelligence does not do that job better btw. It doesn’t do the indexing at all. It just wastes energy. And, as another commenter pointed out, it actively makes the situation you describe worse.
I may not be quite as cynical about AI in its present state, as I feel it has helped me find solutions that had been evading me. I do think it is massively overhyped, however, and am concerned about the directions it's heading. Concerned a lot, actually.
fediverse is it’s trying to layer some order over a decentralized physical infrastructure.
I'm not sure I follow you. The Fediverse is about social networking which is not all, not even most of the internet.
And the internet itself is a layer of order over its decentralised physical infrastructure.
And the internet itself is a layer of order over its decentralised physical infrastructure.
That's fair. But you can have different layers of order that manifest in various human pursuits, and the Fediverse seems a fertile ground for exploring such ideas.
In the early days, the Internet was just a bunch of institutional networks all linked together, with little islands of data that may have had some internal organization. But no big picture view existed, and it never really emerged.
It seems to me that what's happened in recent years is individual companies that offer some sort of information service have been centralizing and gatekeeping their product more and more, which does bring some order but at the expense of what makes a decentralized network powerful in the first place.
It seems to me that what’s happened in recent years is individual companies that offer some sort of information service have been centralizing and gatekeeping
I thought that's what you meant but your statement that the fediverse "trying to layer some order over a decentralized physical infrastructure" is totally removed from that.
BTW I'm not cynical about AI, I'm against it. There's a difference.
Were Luddites "cynical" about mechanical weaving? I don't think so, but the people who called them Luddites in the first place certainly were.
I do some freelancing that involves online desk research, and for some topics, it's like pulling teeth to find a source written by a real human. Even non-Google search engines can't help but give me dozens of obviously AI-written pages. Which is why I as a human am getting paid to do real research. It's literally wading through slop to find anything a human did.
The thing is, AI companies are terrified of model collapse because if they train AI systems on AI output, the model becomes useless. But the rate at which websites on any topic are just slop being scraped to train a new model seems to be either unsustainable for using the open internet for that training, or an assured way to reach model collapase quickly.
There are different theories around model collapse, so there might be a way to avoid it even if most of your training data is AI generated. The jury is still out though
The thing is, AI companies are terrified of model collapse because if they train AI systems on AI output, the model becomes useless.
That's actually not a fact, more of a possibility. As far as am I aware, reinforcement learning from AI feedback (RLAIF) is currently employed by some of the most capable frontier AI companies; Anthropic openly admits to it, and their models are arguably the best.
A term that is often thrown around nowadays is recursive self-improvement (RSI), and it seems that there are at this time more knowledgeable proponents of RSI for intelligence development than opponents who raise the risk of model collapse.
But yes, training on today's sloppified public internet would probably only have downsides.
That is just so depressing. It's eating its own garbage to create more content. Like some image that's been run through multiple passes of lossy compression with different algorithms.
In the early days of the internet there was Gopher which had a strong hierarchy for sorting and returning docs: https://en.wikipedia.org/wiki/Gopher_(protocol)
We could resurrect it, but sadly the AI Slop will still be there I imagine.