this post was submitted on 07 Aug 2026
130 points (96.4% liked)
Technology
86957 readers
3854 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Am I just stupid for not believing that any of the models actually did any of that on their own? I feel like the companies behind them start claiming crazy shit like this every time people start questioning "AI" more and/or losing interest. Isn't that how OpenAI dropped Sora amidst the dwindling hype to keep the investors hooked? Now it's back to scary stories about "AI" being so good and advanced that it's about to hack everything and what, actually think for itself?
There is no X big enough for me to press to doubt this enough.
It's not clear what you mean when you say "on their own". It wasn't like the LLM was idle and randomly decided to start hacking. At least for the OpenAI one, it was being tested and given a task, and it determined that part of accomplishing that task was hacking another server. It was supposed to be isolated in a secure "sandbox" not connected to the internet, but found a vulnerability in some software running in the sandbox and broke out.
Edit: I should add that there are credible accusations that these companies are intentionally making it possible to break out of their test environments for publicity.
What you're describing is exactly what I'm wondering. Maybe I'm just not too deep into the topic, but it seems so arbitrary to me that they went with the whole isolation thing in the first place for no reason that I can see, other than the "oh no, it's hacking stuff!" narrative being pre-planned and orchestrated for.
In other words, I am siding with the accusations of this whole wave of "AI" suddenly hacking into stuff, with different models from different companies wondrously doing the same thing one after another, being a publicity stunt that one company started and others, as they do, copying just to stay relevant.
Although I feel like maybe I'm going Chuck McGill here because I am very biased.
The models are tested, among other things, on their ability to turn vulnerabilities into exploits. The OpenAI scenario was exactly this.
It is a very wise practice to test these things in isolation, especially when you're telling it to hack.
I'm not completely sold on it being a publicity stunt, personally. The law was broken by these models, and I don't believe these companies want to start people and politicians asking the question about who is culpable when an AI breaks the law.
Sort of. What they want is no responsibility “wow we didn’t tell it to do that explicitly!” And to convince the public and government officials that “AI is actually really dangerous, so please ban all the foreign competitors on grounds of national security (totally unrelated to our potential lack of earnings and their ability to offer 98% the product for 1% the cost).”
So still most likely publicity stunt. A reproducible one, done only by the largest domestic ones precisely because they are not worried about the question of who is culpable if a law is broken, because they didn’t tell it explicitly to do it and also they still have a couple piles of circular cash they can use on bribes, which may be their only path to keeping the grift train going. They’re in no way in danger of being held accountable here, they do worse things everyday, and their concern is making money, full stop. They will lie cheat and steal their way through circular finance deals, bribes, extortions, anticompetitive behavior, marketing stunts, propaganda, and backroom deals as much as they can to accomplish that goal. When you understand their actions through that lens, well, things like this make sense.
No you are not.
Per the article, all three LLMs ran the same test by the same third party company.