jj4211

joined 3 years ago
[–] jj4211@lemmy.world 1 points 3 hours ago

The content on Internet will invariably biased towards the novelty. Between that bias and enormous marketing spin and the most self important people gravitating towards it, it's an expected reality. One that will be applicable to some scenarios.

Content saying that for some situations, the existing methods remain best isn't going to light the world on fire. Also, tech folks tend to be more shy about "not getting" a seemingly great new tech.

In terms of how much it applies to an individual situation, it's too nuanced to really make a universal judgement. But in my local circle where I can grasp the nuance, the teams that have gone all in on deeply attention are teams that already were kind of crappy.

[–] jj4211@lemmy.world 4 points 5 hours ago (1 children)

I have both used agents and have been the "victim" of heavy agent users.

If you can provide an utterly well bounded problem with perfectly verifiable criteria for it to retry against until the tests pass and the tests are full and valid for the use case, it can work. Similarly, if your code's reality is forgiving and flexible, and a 'close enough' result is good enough to get the human software user in the right ball park, you might be able to extract decent behavior.

However, if there is a means by which the answer can 'look correct' yet be wrong in the real world and the real world scenarios require accuracy and precision, there's huge gaps.

Especially if the agents control the coding and the test case generation, seen plenty of times where it talked itself out of a test case that was failing when the test case was in fact correctly showing a flaw.

[–] jj4211@lemmy.world 2 points 5 hours ago (2 children)

A number of people have shared their experience with you and you seem to be unwilling to acknowledge that your assertion of universal slopping it up as reality is not utterly universal.

[–] jj4211@lemmy.world 2 points 5 hours ago (1 children)

Frankly, your use case sounds like something has gone horribly wrong from the outset. If something needs several thousand regex patterns, then something is very very wrong, and there's zero chance you have a comprehensive solution as it stands. Either trying to use regex for a use case that some natively AI approach might actually be warranted for, or regex is being used poorly and each one is too limited, or some other approach entirely is warranted.

[–] jj4211@lemmy.world 1 points 5 hours ago (5 children)

Perhaps your experience is the exception?

[–] jj4211@lemmy.world 1 points 5 hours ago (1 children)

The problem is that every time, in the moment, people advocate and then when it comes up short, come back with "Oh, you did XYZ 3.2? Yeah that's busted, 3.3 can really do it" and rinse and repeat and it's hard to take that argument credibly when it's a treadmill of dissing yesterday's tech but swearing today's is different.

The current best of breed I have very limited exposure to (too rich for my blood), but it didn't seem overwhelmingly a slam dunk in my interaction. Incremental value going from more modest models to those don't seem to justify the price tag.

And every developer that is merely curating fully agentic workflow I have dealt with has pretty shit functionality and no idea what the hell they are doing. Whatever the rhetoric is, the results are just utter shit. Somewhat better if the problem domain can be infinitely retried without consequence with results that can be perfectly verified to let it drive eternal retries (while burning through token budget), but generally I just see pretty shit software.

On the other hand, a lot of these groups that I say has pretty shit AI software formerly had pretty shit normal software. Problem being clueless management now thinking they must be smart because they say things more aligned to the AI hype.

[–] jj4211@lemmy.world 10 points 5 hours ago (3 children)

Agents can already write much more organized code than any developer in my company.

Then you must have nothing but absolute shit developers.

[–] jj4211@lemmy.world 3 points 5 hours ago

Been suffering the hacks and downtimes for over two decades waiting for the day when corpos come around on this.

Hard to continue holding out hope.

[–] jj4211@lemmy.world 4 points 6 hours ago

The big problem is that sometimes it's "good" at diagnosing issues and sometimes it's very bad and in both cases it looks the same if you don't have a way to know yourself.

I've seen some of the fodder people consider 'well formed playbooks' and it's also a bit of a crapshoot there. It can speed up some tedium, but so many users get in over their heads.

Today had a junior proudly provide an untested playbook that just did everything wrong and even if it did what it purported to do, it would have been hard coded to certain things that are not what we want to do.

[–] jj4211@lemmy.world 13 points 6 hours ago (1 children)

My god, had someone send in a merge request with utter garbage, with Claude describing some nonsensical "root cause" as to why a function didn't work. Claude made two changes:

  • If the current scheme failed, use an entirely different scheme it claimed the software supports that does not work and was never documented as a way to work.
  • If that scheme fails, swallow the error and don't report the problem, leaving the problem silently in place but out of the user's view at least for a short while. This is actually the change that 'fixed' the problem.

So I needed to know the debug data for what really went wrong, but the submitter got all pissy and said Claude already sorted it out and why can't I just accept their slop merge request as-is, refusing to believe that all it did was catch and silently drop any and all errors associated with the code instead of fixing it. Wouldn't run the procedure I stated would show them they had messed up data as a result of their 'fix'.

[–] jj4211@lemmy.world 3 points 6 hours ago

Funny, but actually a bit on the nose, repeatedly in the show.

Phaser blast? In short order they have shields that block them. Time for some good old punching, which somehow the drones always seemed to be fairly vulnerable to. Also holodeck tommy gun made short work of them. Seems like starfleet should have just issued good old fashioned guns...

[–] jj4211@lemmy.world 1 points 6 hours ago

Dear god bluetooth is such an ugly reality that seems like it could have been sooo much better.

view more: next ›