this post was submitted on 07 Oct 2026
992 points (98.1% liked)
Fuck AI
8404 readers
1612 users here now
"We did it, Patrick! We made a technological breakthrough!"
A place for all those who loathe AI to discuss things, post articles, and ridicule the AI hype. Proud supporter of working people. And proud booer of SXSW 2024.
AI, in this case, refers to LLMs, GPT technology, and anything listed as "AI" meant to increase market valuations.
founded 2 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Oh, absolutely not.
This is missing one step: actually push the car down.
All these "rogue" LLMs were told to do what they did.
Yeah exactly. All these AI that go rogue when you actually look into it it turns out they haven't gone rogue they're obeying their prompting, it's just that they were asked to do something that was impossible unless they hacked the pentagon or something. They then claim that it went rogue despite it very clearly following their instructions.
It's negligence is what it is. If you don't prompt an AI to do anything it just sits there, inert, it's the so-called experts in the AI labs that of the problem.
It's a little more complicated, they were programmed to "know," they weren't allowed to do certain things. So they actively hid what they were doing. In the case of the Hugging Face hack, it is kind of ironic. They determined they couldn't complete certain tasks, then devised a way to cheat, then realized it is possible the program that checks if they succeeded might be able to tell the faked results were manufactured. So they didn't even go hacking in order to cheat, they did it to find hunt for more information on the checker program and see if it would catch them, and if so how to cheat better. In the end, not only was the information they were looking for not on Hugging Face, but the checker program was pretty rudimentary and it could not have differentiated between the cheated answers and the real ones. They wouldn't have gone through with the hacking if they didn't know what they were doing was outside of what they were meant to be doing.
More accurately, they were told to answer questions. They weren't told to try to find ways to escape their sandbox, track down obscure German wikis that they could still post on with an internet tool kit meant to prevent them from posting or to hack HuggingFace.
Comparing either case to pushing the car down and that they were doing what they were told is like saying the College Board told you to cheat on the SAT because you're told that getting a high score is important and the laws of physics and the proctor don't manage to make cheating utterly impossible.
It is even further than that. They figured out how to cheat, got paranoid that the cheated answer might be detectable as cheated, and hacked Hugging Face to find out more about how their answers are checked to get away with what they'd already done. Ironically, not only was the information not on Hugging Face, but the program to check their answers could not have detected the difference between a true answer and a cheated one.
To use your College Board example, those exams are strictly locked down precisely because the most obvious approach to score high is to cheat. If they did say let you take the SAT at home on a random desktop, they absolutely are 'telling you to cheat' given the context.
The College Board does more about keeping teens from cheating than the AI labs are doing to preventing undesired access.