Mildly Infuriating
Home to all things "Mildly Infuriating" Not infuriating, not enraging. Mildly Infuriating. All posts should reflect that. Please post actually infuriating posts to !actually_infuriating@lemmy.world
I want my day mildly ruined, not completely ruined. Please remember to refrain from reposting old content. If you post a post from reddit it is good practice to include a link and credit the OP. I'm not about stealing content!
It's just good to get something in this website for casual viewing whilst refreshing original content is added overtime.
Rules:
1. Be Respectful
Refrain from using harmful language pertaining to a protected characteristic: e.g. race, gender, sexuality, disability or religion.
Refrain from being argumentative when responding or commenting to posts/replies. Personal attacks are not welcome here.
...
2. No Illegal Content
Content that violates the law. Any post/comment found to be in breach of common law will be removed and given to the authorities if required.
That means: -No promoting violence/threats against any individuals
-No CSA content or Revenge Porn
-No sharing private/personal information (Doxxing)
...
3. No Spam
Posting the same post, no matter the intent is against the rules.
-If you have posted content, please refrain from re-posting said content within this community.
-Do not spam posts with intent to harass, annoy, bully, advertise, scam or harm this community.
-No posting Scams/Advertisements/Phishing Links/IP Grabbers
-No Bots, Bots will be banned from the community.
...
4. No Porn/Explicit
Content
-Do not post explicit content. Lemmy.World is not the instance for NSFW content.
-Do not post Gore or Shock Content.
...
5. No Enciting Harassment,
Brigading, Doxxing or Witch Hunts
-Do not Brigade other Communities
-No calls to action against other communities/users within Lemmy or outside of Lemmy.
-No Witch Hunts against users/communities.
-No content that harasses members within or outside of the community.
...
6. NSFW should be behind NSFW tags.
-Content that is NSFW should be behind NSFW tags.
-Content that might be distressing should be kept behind NSFW tags.
...
7. Content should match the theme of this community.
-Content should be Mildly infuriating. If your post better fits !Actually_Infuriating put it there.
-The Community !actuallyinfuriating has been born so that's where you should post the big stuff.
...
8. Reposting of Reddit content is permitted, but attribution is not required in any way. No links to Reddit in post body
-If you would like to provide a source link, do so in the comments but not in the post body.
...
...
Also check out:
Partnered Communities:
Reach out to LillianVS for inclusion on the sidebar.
All communities included on the sidebar are to be made in compliance with the instance rules.
view the rest of the comments
I'm curious what you are using. The free versions of chatgpt have been like that for me, but even Gemini flash with extended thinking, also free for a while longer, is giving me pretty reliable results as long as there training data out there to derive an answer from. The higher (paid) Claude models will one shot most coding tasks.
I mean as long as the coding tasks are simple, and seeded with enough context, then sure. As in "create tinder clone works", but I still have to check the output in my company codebase. We are leveraging llms a lot for code writing, full agentic pipelines, loops, all the shiny new approaches.
The outcomes in all cases are still mediocre or below.
We have twice as many open prod bugs as a year ago.
I've not yet fucked with Claude. I don't want to pay for it, and I really don't like the surveillance aspect of these centralized systems. Mostly I'm using Gemini, whatever DDG had in their search results, and local models I've been fiddling with (like Qwen3.8 right now).
All more or less garbage once I get into the details of anything on the edge of my expertise.
I got a free year of perplexity.ai which gives me limited access to claude sonnet and yeah, its honestly pretty good.
I've been using OpenCode with whateverthefuck free models they have listed on there and they all seem to do fine with agentic tasks like building me scripts or executables to make my work tasks easier.
I used Gemini at the start with "Frontier Knowledge" and it seemed to do worse than the ones listed on OpenCode, but maybe that's because i could only do like three prompts a week since I refuse to pay into an AI.
end of the day, its just LLMs writing code for me, but I cannot see how this would be useful for a large scale code base, but also #NotAProgrammer.
Are you using Gemini in flash extended thinking? (Not flash light) . I haven't had many hallucinations other than cases where the training data it needs just doesn't exist (cases where I can't find the answers by googling either)
Claude can one shot tasks until you get a larger system then it completely shits itself. These models are nothing more than autocomplete, and they can’t hold large systems in their heads. Anthropic literally tried to rewrite all of bun using Claude, they said they did it and yet it still hasn’t released six months later.
This has gotten a lot better for me by having a "send out scouts" skill that has a lower tier model agent search through the codebase before it starts to plan. Has handled my companies giant monolith pretty well and even can handle cross repo features as well.
Claude code has been using the new rust bun for ~5 months now and has been working fine, and bun 1.4 that released in August is using the rust rewrite
It is worth noting that the repo had multiple Anthropic employees commiting to it after the bun rewrite and before release, and that so far the only successful use cases that Anthropic or OpenAi showed are... rewrites from one language to another given sufficiently large test suites.
Yeah I certainly wouldn't advocate building a whole business around code it wrote. But for small personal tasks it hasn't let me down. Building custom server applications, desktop applications, Firefox add-ons, upgrading my homeassistant 10 versions over a couple weeks without letting anything break. These sort of things it handles pretty easily and are all things I wouldnt get done without it.
How do you verify that everything it tells you is correct?
You check the links/references it gives you.
Gemini does a pretty good job of this because it doesn't seem to have much built-in knowledge. Instead, it just searches the Internet on your behalf and returns summarized results with links to where it got that specific information.
I use it to search for scientific research all the time and the summaries often aren't detailed enough so I actually click on those links. I've yet to encounter a situation where it fucked that up (invented links that don't exist) but I have heard about it happening.
So far, the summaries have seemed to be pretty spot-on when it comes to biology papers 🤷
By writing instructions to insist that it double verifies every (non obvious) claim with a minimum of two independent sources. I also told mine to always assume that the initial prompt is missing crucial context, and to ask as many follow-up questions as necessary until it has enough information to provide the answer to the question I'm really asking. (For speed and efficiency you can even make it give you multiple choice options to click on.) Because sometimes the problem isn't with the LLM, but with the user asking the wrong questions.
Using those two instructions alone, I've encountered considerably fewer hallucinations, and when I'm still not certain, I can simply click on the sources linked next to every single claim the AI makes.
Right, but that sounds like you implicitly trust the LLM and only verify statements when they seem off. So your intuition is the final arbiter if truth?
No I'm saying that I don't have to trust it because I instructed it to put two independent sources next to every claim it makes. I just check the sources.
It definitely has an element of garbage in, garbage out.