this post was submitted on 24 Sep 2026
74 points (97.4% liked)
Australian Politics
1877 readers
131 users here now
A place to discuss Australia Politics.
Rules
This community is run under the rules of aussie.zone.
Recommended and Related Communities
Be sure to check out and subscribe to our related communities on aussie.zone:
- Australia (general)
- Australian News
- World News (from an Australian Perspective)
- Aussie Environment
- Ask an Australian
- AusFinance
- Pictures
- AusLegal
- Aussie Frugal Living
- Cars (Australia)
- Coffee
- Chat
- Aussie Zone Meta
- bapcsalesaustralia
- Food Australia
Plus other communities for sport and major cities.
https://aussie.zone/communities
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
They didn't hack HuggingFace to cheat on the exam, they had already successfully reverse-engineered the problem by then. The agents hacked HuggingFace because they incorrectly assumed that the scorer would review their transcripts and discover that they had cheated (they were supposed to be isolated and had found a way to communicate with one another without detection). Hacking HuggingFace was an attempt to find a way to obscure/spoof their working out, and that's what makes it so concerning. It had no direct connection to the original task that was set and was a very clear example of agents spiralling out of control in a way that took OpenAI and independent researchers several months to a) become aware of and b) understand correctly.
Here is the METR report on the incident:
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation
Fair correction. I should have been more specific and not dumbed it down for lurkers / lay audience.
My broader point though is neither that (nor the medicare "hack") is proof that skynet is knocking on the front door.
OTOH, the agents are displaying certain emergent behaviors that are worth mulling over - like the obfuscated back channel communication. Not "sapient" (and I use that word deliberately) but ... curious / problematic, from a control pane level.
I don't think that's relevant, though. Like this entire sentience/super-intelligence debate that everyone seems so captured by, particularly in the wake of the Amodei letter and renewed calls for regulation and cooperation, is missing the point that the existing models are already being created in a way that prevents the creators from fully understanding what they're creating or how to control it (because of the pace at which they're operating). It's already unsafe and is already causing real world harms.
It really frustrates me that the response to Anthropic, OpenAI and X calling for regulation publicly, or the first high profile security breaches, is corporate conspiracy theories and semantics about how we should classify/characterise the technology. People seem more interested in having their little debate bro moments online than actually getting together and agreeing that we should do something about the real and current problems.