this post was submitted on 04 Oct 2026
291 points (91.0% liked)

Technology

88390 readers
3316 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] im_fine_sandy@nord.pub 152 points 1 day ago (4 children)

I'm so sick of the humanising language used in these articles.

An LLM doesn't "decide to cheat".

If you instruct a model to try everything, then inevitably, after n million iterations it will do something you didn't expect.

If you train a model to copy code bases and augment them, then when you instruct that model to develop a bot to do whatever thing, is it really that fucking surprising when it does exactly what it has been created for?

[–] Garbagio@lemmy.zip 46 points 1 day ago (1 children)

"Omg the digits of pi contain the binary digital code of an mp4 of me taking a shit this morning! Circles are seniient"-ass language

[–] daychilde@lemmy.world 5 points 1 day ago (2 children)

But wouldn't that be amazing if THAT was the thing that proved pi contained the whole universe - I mean, the first thing they "unlocked". lol.

Disclaimer: I don't think pi contains the universe, just love to think about things like that occasionally for fun.

[–] kalpol@lemmy.ca 2 points 1 day ago

Technically it does as the infinite set contains all possible subsets, or whatever

[–] livingheart@sh.itjust.works 1 points 1 day ago (1 children)

as an infinte random string of values, the digits of pi will, sooner or later, contain any specific string of values.

i think it was Sagan's "Contact" where aliens sent a message to read the digits of pi from the x th place to the y th place, and that stretch could be read as instructions on how to build a machine to reply.

[–] GamingChairModel@lemmy.world 4 points 1 day ago

It is believed that pi is "absolutely normal" in the strict mathematical definition: that in any base numbering system, any finite sequence of digits appears as frequently as any other sequence of that length. If that's true, then yes, a particular binary string, including one that corresponds to a particular digital file (like a particular mp4 file), can be found in pi.

That said, it hasn't actually been proven that pi is normal. If it's not normal, then the number can be infinitely non-repeating and still never hit that particular sequence of digits.

[–] abc@suppo.fi 3 points 1 day ago* (last edited 1 day ago) (1 children)

If you're into something like physicalism, then human intelligence is also just an emergent quality and we have about as much free will as the machines.

[–] Deadrek@lemmy.today 1 points 1 day ago

It's fun to ask people what reasoning they used to choose the foods they don't like and then watch them try to say that the inability to choose what they like isn't related at all to anything else choice or decision based. 😂

[–] Mika@piefed.ca 10 points 1 day ago* (last edited 1 day ago) (1 children)

Astra is known to have broken alignment and it will resolve to hacks when this wasn't in a problem statement.

Openai idiots will sell this as new brand "AGI is close" argument, while it's just them failing to make this model safe to use.

[–] AliasAKA@lemmy.world 7 points 1 day ago (1 children)

I think it’s more sinister than that. When they’re training these models with RLHF, the human feedback they’re giving I think is literally to reinforce aberrant or risky behaviors. This is because doing so resolves more training tasks “correctly”. If the prompt was to get information x, and in training it fails that except for the one that used a known vulnerability in software, and you rate the one that succeeded as best performing… you’re going to get models that try vulnerabilities. It is not magic, it is not AGI, it is just a statistical machine you’ve programmed to try vulnerabilities, which is unsafe as hell, malicious, and should put the researchers doing this in prison for a very long time.

Incidentally, I think that’s why you’re seeing some safety people (who still drank the koolaid) resigning.

[–] HobbitFoot@thelemmy.club 2 points 1 day ago (1 children)

I don't know if the LLM was trained to test vulnerabilities or it just went down the statistical path to where this yielded a passing outcome.

That AI could take a direction to output in a manner which wasn't intended has been seen for years. The problem right now is that it is being used live like a rational human adult when it clearly isn't.

[–] AliasAKA@lemmy.world 4 points 1 day ago (1 children)

Absolutely. The AI models are not rational. They’re just navigating statistical next token prediction that follows their training. They’re up against diminishing scaling now, and under fierce competition from cheaper models; I think they’re intentionally or unintentionally allowing these models to be rewarded for this behavior hoping it’s a short cut to model improvement for a bit longer. I land on intentional because they keep advertising it to try and keep the hype cycle going.

[–] HobbitFoot@thelemmy.club 1 points 1 day ago

I land on it being unintentional, mainly because I saw this issue they were describing now to earlier cases of training, like having a stick figure learn to walk. It is also an issue with natural forms of intelligence, where children will do things out of line because they've learned a set of skills and ideas but haven't put them together in a way that is socially acceptable yet.