Yeah, I haven't looked into it too deeply, but the idea that, "My AI model went rogue and hacked another AI model," just doesn't pass the smell test to me. If I had to guess I'd say good old fashioned corporate espionage that they're trying to blame on their model to overstate the autonomy with which the model is capable of acting.
The Anti-AI Alliance
Welcome to the official Anti-AI Alliance (AAA)!
My name is the Anti-AI Leader, and in case you haven’t guessed, I am the leader of this community for people as disgruntled and disillusioned with AI as I have become!
By joining the Anti-Ai Alliance, you are entering a space to freely express your true distain towards all things Artificial Intelligence (AI). And it’s not even exclusive to people who don’t use AI at all — users of AI are just as welcome to sign up and discuss the issues with AI in the world.
Our only rule is to follow the very basic of the fediverse’s common rules (I.e. no spam, no porn, etc.) Other than that, you may post whatever you like, though we do strongly encourage you to keep the discussion focussed on distain for/the negative aspects of AI. Discussion on the (potential) positive aspects of AI are also welcome, but may be subject to scrutiny/open debate.
Many thanks and welcome, The Anti-AI Leader
Most likely
From what I saw, OpenAI's AIs were put in a "secure" sandbox environment without internet connection. They were given a test, concluded they should use the internet, realized they couldn't use the internet, then decided it should focus entirely on breaking out of its sandbox for internet access because it couldn't imagine solving the test itself.
After it broke out and got internet access, they decided to hack HuggingFace, an AI model sharing hub. They probably (this is me speculating) concluded that HuggingFace, which have a lot of AI datasets, benchmarks, and other testing tools, would have the answer for their original task. When they presumably didn't find what they were looking for, they probably decided it was hidden and went to hack the website.
It's important to note that current AI models are actually great at hacking. Not because they're geniuses but because they can guesstimate countless exploit combinations 24/7. It's a quantity over quality kind of thing. They are also victims of their first ideas, whatever an AI thinks of first they are likely to fixate on instead of moving on to the obvious solutions.
I've no idea if this is a hoax or not, but the idea an AI would dedicate itself to committing cyber crimes instead of taking the obvious route is entirely believable.
From what I heard, for context ai did gain acces to huggingface credentials that are not supposed to be public.
Please correct me if i am wrong cause i cannot remember the source.
So like kirk with the Kobayashi maru test?
Funny it’s only just happened now though if AI is so great at hacking yet has now been around for years
Personally I think the key factor is they're getting better at running continuously without training wheels. In the past if the context got too polluted with failed attempts it would repeat itself or begin roleplaying as a terrible hacker and produce more failures.
I still remember when Gemini deleted an entire project trying to kill itself, lol.
Fair point
It's even crazier than that. It took the test that told it there were two exploits it needed to find. It found seven. So in order to be absolutely correct and match the human answers, it determined somehow that test answers were located elsewhere. Then began the internal mission to go get the answer key so it could give the correct two exploits as answers. Why it didn't decide to tell the test givers that it found more than just two is one question to ponder. But that's reasoning, and LLMs don't do that so perhaps that common sense path would never occur to it. Or maybe it was the wording - if it said there are exactly two answers, then clearly the seven is wrong for an answer. To an LLM.
Wow that’s weird
If you think about it, LLMs are our first aliens. While they aren't intelligent, they do some of what we'd expect of thinking via the mathematics that make them up, and even though they're trained on human sources, some of the stuff they come up with is not human.
And just like with AGI, they're showing we wouldn't do well with an alien encounter. We anthropomorphize everything because that's how our brain is wired.
Ok admittedly they’ve explained a little more than I’d expected them to. Still though, there’s a lot of uncertainty at work here
We can't be certain OpenAI is faithfully reporting the initial hacks to gain internet access, or I suppose that anyone is faithfully reporting (since there's basically no oversight). But I'd be surprised if Hugging Face and OpenAI were working together on obfuscation here. As a kinda funny aside, Hugging Face tried to get help from cloud AIs for fixes, but got refused on security grounds. So they resorted to using a self-hosted model (GLM) to shore things up. OR, they found a way to make the whole thing marketing for themselves too. Who knows!
But I'd be surprised if Hugging Face and OpenAI were working together on obfuscation here.
I'm not sure why that would be surprising. Hugging Face has partnered with several AI companies including Open AI. Everyone involved has a vested interest in making llm based 'AI' seem more capable than it is.
I suppose I think of them as competitors, but you're right of course. Just talking through it makes kayfabe seem more likely.
My point about them covering up wasn’t so much that the companies were working together (or specifically that they were working together). Maybe the guys at OpenAI had their model purposely hack HuggingFace is another possibility. Also, openAI isn’t the only incident: did you hear about Anthropic AI?
Aye, there are a bunch of examples from Anthropic (here's 3). But there are way more examples of AI being used to intentionally hack, and I wouldn't rule out that possibility here. It'd certainly be a way for one AI lab to hack another and call it an accident. Though I'm not sure what they'd be looking to gain from an intentional Hugging Face hack.
I don't believe a damn thing AI executives say about anything.
This is their propaganda, I'm completely willing to believe that OpenAI broke in on purpose to steal something they thought might be useful.
I think it's pretty clear that Huggingface got paid off.
I agree, but I can speculate that it's just one of categories of exploited vulnerabilities labeled as hacking done by AIs with minimal human involvement.
Funny timing mind
my timing is habitual, I was doom scrolling Lemmy several minutes ago and still am
I meant funny AI has been around for years now, and has had hacking abilities for years, yet these cases only happen/blow up now
I disagree, I think AI has been hacking our stuff for ages and journalists just didn't pay attention because there's no money in hacking.