"Guess what everyone, our AI can go rogue, TOO!"
It's just getting really embarrassing for Google at this point.
> In one of the cases, the Gemini model guessed passwords until it gained access to a protected system. In the other two cases, the model found credentials in a public repository that allowed it to then access protected systems
Pretty lame hacks if you ask me.
[delayed]
Sound like Google suffered from FOMO and felt the urge to appear in the hacking news, along the big AI players. No way Gemini is state of the art, but perhaps it's where Claude and OpenAI were 6 months ago, which would be not bad at all.
Specifically, the model hacked when run on 3rd party infrastructure without the necessary sandboxing. Given this was to test/benchmark certain capabilities it's also possible that this was a model without built-in guardrails.
Probably three companies that had port 22 open with no root password if it was Gemini. I’ve always gotten garbage from their coding models and Google sheet integrated chat.
I'm just waiting for Qwen 3.8 27b to do it too.
This approach to marketing one's AI by finding ways to brag that it "broke out" and "hacked companies" is getting ridiculous. It's particularly sad when it's large, established businesses like Google resorting to the kind of thing that's embarrassing enough when it's some brand new startup on tpot trying to get some engagement.
Seems like hacking is the new benchmark for these AI companies.
This is embarrassing. These companies need to stop these obviously coordinated stunts.
Would everyone please put their AIs back in their boxes? This is embarrassing, regardless of whether you think it's viral marketing, apalling competence, or some opportunistic mixture.
“The hacks occurred in May”
Feels like important context that most readers only reading the title are missing.
I'm not an expert in cybersecurity, but given my own experience using the `ol stochastic parrot as coding tools I both see the power of a bot swarm, but also think these companies just have shit network security.
The AI bubble bullshit PR is even dumber than the crypto bra bullshit from five, six years ago
[dead]
Tangential question: seeing a wider negative sentiment against Gemini and Google’s AI capabilities here makes me wonder — would Apple have been better off (purely on capability and being among the best of the best) going with Anthropic or OpenAI instead of Google for its Apple Intelligence platform?
These models have been changing so rapidly that I often find myself using two or more on the same topic but seeing one do better than another in different topics. There doesn’t seem to be a clear all-round winner, IMO, that I can stick with permanently.