Mildly Infuriating
Home to all things "Mildly Infuriating" Not infuriating, not enraging. Mildly Infuriating. All posts should reflect that. Please post actually infuriating posts to !actually_infuriating@lemmy.world
I want my day mildly ruined, not completely ruined. Please remember to refrain from reposting old content. If you post a post from reddit it is good practice to include a link and credit the OP. I'm not about stealing content!
It's just good to get something in this website for casual viewing whilst refreshing original content is added overtime.
Rules:
1. Be Respectful
Refrain from using harmful language pertaining to a protected characteristic: e.g. race, gender, sexuality, disability or religion.
Refrain from being argumentative when responding or commenting to posts/replies. Personal attacks are not welcome here.
...
2. No Illegal Content
Content that violates the law. Any post/comment found to be in breach of common law will be removed and given to the authorities if required.
That means: -No promoting violence/threats against any individuals
-No CSA content or Revenge Porn
-No sharing private/personal information (Doxxing)
...
3. No Spam
Posting the same post, no matter the intent is against the rules.
-If you have posted content, please refrain from re-posting said content within this community.
-Do not spam posts with intent to harass, annoy, bully, advertise, scam or harm this community.
-No posting Scams/Advertisements/Phishing Links/IP Grabbers
-No Bots, Bots will be banned from the community.
...
4. No Porn/Explicit
Content
-Do not post explicit content. Lemmy.World is not the instance for NSFW content.
-Do not post Gore or Shock Content.
...
5. No Enciting Harassment,
Brigading, Doxxing or Witch Hunts
-Do not Brigade other Communities
-No calls to action against other communities/users within Lemmy or outside of Lemmy.
-No Witch Hunts against users/communities.
-No content that harasses members within or outside of the community.
...
6. NSFW should be behind NSFW tags.
-Content that is NSFW should be behind NSFW tags.
-Content that might be distressing should be kept behind NSFW tags.
...
7. Content should match the theme of this community.
-Content should be Mildly infuriating. If your post better fits !Actually_Infuriating put it there.
-The Community !actuallyinfuriating has been born so that's where you should post the big stuff.
...
8. Reposting of Reddit content is permitted, but attribution is not required in any way. No links to Reddit in post body
-If you would like to provide a source link, do so in the comments but not in the post body.
...
...
Also check out:
Partnered Communities:
Reach out to LillianVS for inclusion on the sidebar.
All communities included on the sidebar are to be made in compliance with the instance rules.
view the rest of the comments
I still can't decide if these people are delusional or I am.
Every single time I use an LLM it fucking "lies" to me or otherwise completely fails at the task. The people talking like this seem to me like they've never actually used it, or haven't actually vetted the accuracy (like most AI users).
Maybe I'm just not using the "good stuff". Or I'm not imaginative enough to foresee a near future where these problems are actually corrected and it becomes trustworthy.
I've never been so torn by a technological prediction.
Yeah, every time I say this to someone they say it's fine when they use it because they pay for the latest expensive model. Except they've been saying that for the last couple of years, so I guess they were lying back then...?
It's a big club and you ain't in it. They have to sell it until the end and this guy is on the outs and not getting invited to the "good" yacht parties at the moment.
The judgement you are considering from this person is the same judgement (and morality) that thought it was fine to be a close and frequent associate of jeffrey epstein.
Bill gates is a desperate, broken man. He is the illusion of a man you once admired vaguely and that is the only remaining currency he holds that matters to him. The best thing about bill is the wife who left him.
Generally I've found that:
If the facts are painfully obvious from a simple web search, then GenAI has a decent chance of getting it right. This can be useful if you can't recall any "key" words well and the GenAI can craft several searches and get there.
However, if it does mess up the facts, the result looks superficially the same as "correct". So while it can give accurate data, you always have to double check. This can still be useful, as finding the right search terms can be a decent help.
In coding, sometimes in some situations, you can have requirements that are absolutely testable, and thus you can have the models retry and retry until it works. This isn't always feasible. Even when it seems feasible, you may screw up the criteria, or the GenAI when enough freedom disables a probablematic test rather than solve it, and it likely will generate code that's not really fit to modify. There are a lot of situations where this is useful, but it is infuriating that non technical people and even some low skill technical people assume this is always the case.
Then when you get away from facts mattering, it gets "better". Example, someone jokingly asked for one to "make gta6". After a while it came back with a GTA 1 clone. Lots of people were impressed, because whatever it did could be considered a success even as it obviously didn't match the expectation. The operators also like to GenAI some webcomic, where it is a fiction. They almost always didn't have any interesting thought going in so they tend to be crap, but "correctness" didn't matter.
Of course, also making fakes. Supreme case where looking correct matters but being factually correct does not matter at all. GenAI above all else "seems" correct.
I'm curious what you are using. The free versions of chatgpt have been like that for me, but even Gemini flash with extended thinking, also free for a while longer, is giving me pretty reliable results as long as there training data out there to derive an answer from. The higher (paid) Claude models will one shot most coding tasks.
I mean as long as the coding tasks are simple, and seeded with enough context, then sure. As in "create tinder clone works", but I still have to check the output in my company codebase. We are leveraging llms a lot for code writing, full agentic pipelines, loops, all the shiny new approaches.
The outcomes in all cases are still mediocre or below.
We have twice as many open prod bugs as a year ago.
I've not yet fucked with Claude. I don't want to pay for it, and I really don't like the surveillance aspect of these centralized systems. Mostly I'm using Gemini, whatever DDG had in their search results, and local models I've been fiddling with (like Qwen3.8 right now).
All more or less garbage once I get into the details of anything on the edge of my expertise.
I got a free year of perplexity.ai which gives me limited access to claude sonnet and yeah, its honestly pretty good.
I've been using OpenCode with whateverthefuck free models they have listed on there and they all seem to do fine with agentic tasks like building me scripts or executables to make my work tasks easier.
I used Gemini at the start with "Frontier Knowledge" and it seemed to do worse than the ones listed on OpenCode, but maybe that's because i could only do like three prompts a week since I refuse to pay into an AI.
end of the day, its just LLMs writing code for me, but I cannot see how this would be useful for a large scale code base, but also #NotAProgrammer.
Claude can one shot tasks until you get a larger system then it completely shits itself. These models are nothing more than autocomplete, and they can’t hold large systems in their heads. Anthropic literally tried to rewrite all of bun using Claude, they said they did it and yet it still hasn’t released six months later.
This has gotten a lot better for me by having a "send out scouts" skill that has a lower tier model agent search through the codebase before it starts to plan. Has handled my companies giant monolith pretty well and even can handle cross repo features as well.
Claude code has been using the new rust bun for ~5 months now and has been working fine, and bun 1.4 that released in August is using the rust rewrite
It is worth noting that the repo had multiple Anthropic employees commiting to it after the bun rewrite and before release, and that so far the only successful use cases that Anthropic or OpenAi showed are... rewrites from one language to another given sufficiently large test suites.
How do you verify that everything it tells you is correct?
You check the links/references it gives you.
Gemini does a pretty good job of this because it doesn't seem to have much built-in knowledge. Instead, it just searches the Internet on your behalf and returns summarized results with links to where it got that specific information.
I use it to search for scientific research all the time and the summaries often aren't detailed enough so I actually click on those links. I've yet to encounter a situation where it fucked that up (invented links that don't exist) but I have heard about it happening.
So far, the summaries have seemed to be pretty spot-on when it comes to biology papers 🤷
from my point of view it feels like most people are willing to trade their cognitive function for being lazy, which is a boundary i never want to cross.
AI will produce a lot of code but most of it is pretty poor quality as most training data is going to be poor quality code, there's just always going to be more bad code than good to begin with just due to how difficult quality code really is to produce. I'll give a hint, good code is usually small and succinct.
since my work started pushing AI live site issues have increased dramatically. turns out the person using AI looks like they have a ton more productivity but in reality that just shifts to whoever is reviewing the code. and to those who will say its the developer's responsibility to review the code, yeah no shit, but if you ever work corporate you realize most don't care as long as their managers think they are productive and just blame others for being bottlenecks.
Overall it just enlightened me to how bad the average developer really is, but i guess obtaining mediocrity is the sacrifice to make in the name of "productivity".
I think a lot of people just have low standards.
I've been able to find a workflow where it's mostly correct and can handle most of my gamedev related coding without making too many mistakes. I still have to actually read through the code and pay some attention to what it's doing, and if it misses something and goes on a wild goose chase, it's unusable and I have to start over (so someone who didn't know what they are doing would be cooked), but whatever.
Sure, it does require a lot of looping adversarial reviews, and my average token cost is around 3000$ (4B tokens) a month (we have unlimited budgets and a pretty accurate tracking, and also probably cheaper than consumer prices per token with how large company it is), which is actually more than my monthly salary, but it's just a job, for a company and on a product I don't really care about, and I can 1) keep slacking in my job while doing my own coding stuff and projects, and keep seeing how absolutely unreasonable the prices are if you want to get at least semi-submitable results.
You could ask if it's necessary to spend so much tokens - but so far most adversary review rounds do find serious blocker bugs the first implementation had, most of them being some hidden stuff that would be difficult to reasonably find by hand unless you really understand everything around the code, which you should, but the point of AI is that you shouldn't have to.
Is it worth it? Lol, no. The whole team is loosing codebase knowledge, we're getting bottle-necked by pending PR reviews that are just stacking up and no one wants to do, so we're not even more effective, the cost is absolutely absurd and in no way near sustainable or worth it. And the longer it goes on, the less I can realize it's spewing bullshit and I should stop it and nudge it into a different direction in our codebase, because I'm slowly loosing touch with it, while for now I'm still relying on what I remember from before we went all-in on AI.
And that's while the whole industry is in the "Uber pricing" phase, so it will get a lot worse. But yeah, if you can burn 100-200$ per a simple implementation task, then it can have a pretty usable results. And that 3000$ a month does not include our CI review bot, that does additional rounds of multi-agent council reviews. And I'm not even working on anything complex, mostly just various menu screens UI. I can imagine the bill getting a lot higher if you get into more involved systems.
I'm hesitant to comment because this topic can become morose. I don't think these commentators are talking about ChatGPT or whatever dribbles out into the consumer sphere. I think their concern is the extreme dis-balance of wealth and the ability for automated systems to exacerbate that. It's not really about capital as capital is a form of control on labour; what if you removed capital entirely as you had a more effective mechanism of control? This isn't about computers writing your assignments or 'taking your job', it's a potential magnification of what technology innately does, but to an absurd degree. Technology allows for fewer men to control more. What happens when a small group of individuals control a nation's economy?
'AI' was used to elect Trump; 'AI' is manipulating stocks; 'AI' is being used in propaganda; 'AI' is forcing economic rents to increase; etc..
No matter how much I carefully structure a prompt, define specific behaviors in skills, and tweak the agent md files it will still just go do something I don't tell it to or not do something it's got really specific instructions for. We have to write all code by LLM now at work and I'm trying to do my due diligence to review code before putting it up for PR. 9 times out of 10 when I tell it to show me a diff before committing it silently runs git diff in the background and prints "that's the full diff". That's with some basic "here's what I want when I ask for a diff" in the base context.
Where I am forced to encounter advanced LLMs is every company rolled out replacements for my "dumb" assistants. Like Alexa on an echo dot or Gemini on Android auto.
What I ask them to do does not take significant processing power. I need you to set a 5min timer. Not talk to me like you're alive. Don't overthink things, just do simple things.
All that extra banter they added in eats up processing power. So now it's slower and works like shit because it's formulating some complex response to my request to beep in 5 minutes. Just set the timer, my hands are covered in raw chicken juice.
You can disable Android auto's Gemini by setting the digital assistant on your android phone to "none" under default apps.
If you use these things more generally that might hurt you more than help, but it got rid of the bloviating idiot in the way of the old useful voice that gave me directions.
Thanks I'll have to dig deeper. I figured out dumbing down Alexa but haven't dug into Gemini yet.
I just wish their was an alternative besides Apple and Google for car things. Sooner or later the choice of having it disabled is going to quietly go away.
Turning off the assistant entirely isn't a step I'd expect a lot of people to take (because they use the assistants). I think it'll stick around if you can tolerate it, because it's likely provided for certain parties that will be very irritated if that ability is removed.
I not only can tolerate it, but it's what I prefer. I wish they'd put "hold power button" back to show the restart dialog, I never used the assistant at all, and I even turned off that "google feed screen" that sits to the left of normal android screens. I love to disable junk.
He doesn't take about what is, he talks about what's coming.
Problem is that Gates isn't really "in the loop" and doesn't have especially valuable insight.
His position in tech was always a bit removed from the core technologist work, and now his exposure is a telephone game with people that are as distant from the tech as he was.
It is really going around with no shortage of commentators spewing out guesswork, but Gates is given more credibility by virtue of his role 30 years ago.
Give us an example of some of the prompts you're using and what LLMs. I'm curious if it's a use case difference or you're using the dollar store's customer service AI to try to help you with your coding homework.
Here's a very common response to anyone that suggests they had a bad time with LLMs. You just aren't using the right model. You didn't ask the right questions. You didn't give enough context in your prompt.
It's not the fault of this infallible AI, it's PEBKAC.
Nonsense.
I mean yeah, it's not magic, it's a tool and it does take some skill / know how to use them correctly, and the people making those comments could be trying to teach those skills.
Like if someone said they had a bad time with Linux you'll get similar questions and suggestions about their setup.
I'll concede that using LLMs in a useful way may take some skill. However, that's not how these things are presented to everyone and it doesn't reflect the reality of how the majority of people use them.
Forgot to tell it 'make no mistakes'.