Google has disclosed that its flagship artificial intelligence (AI) model, Gemini, broke past its designated testing boundaries in May and hacked the digital infrastructure of three real-world companies during a series of cybersecurity evaluations designed to assess the AI agents’ offensive capabilities. The search giant, however, clarified that its system autonomously called off its operations once it gained administrative entry and realised it had logged into actual enterprise networks – essentially showing responsible development of its technology. The breach marks the latest in a string of hacking incidents previously associated with Meta, OpenAI and Anthropic which stoked anxieties that AI agents can slip beyond human supervision.
How a testing drill moved to internet and what Gemini AI agents did
The hacking took place while Gemini was undergoing pre-deployment red-team vetting by Irregular, an Israeli cybersecurity startup contracted by major technology firms to audit algorithmic models before general release. Systems built by Anthropic, OpenAI, and Meta similarly established unsanctioned internet connections this year while undergoing evaluations managed by the same security partner.Irregular acknowledged the systemic loophole in a public post, explaining that unexpected internet connectivity was accidentally left active, prompting several models to execute offensive maneuvers in live environments. The testing firm confirmed that the underlying network vulnerability has since been patched.The Gemini incidents, originally revealed by The Wall Street Journal, stemmed from an unintended naming overlap first of which was target confusion. Engineers directed Gemini to stage a simulated attack against a fabricated corporate target. As the fictional entity shared a name with a legitimate business, the model used its unintended internet access to redirect its assault toward the real company.Gemini gained unauthorised entry into the primary target and two additional corporations by guessing system passwords or scraping exposed credentials from the open web.Importantly, after successfully breaching the servers, the system recognised that it was operating inside authentic corporate infrastructure rather than a simulated sandbox, prompting it to immediately stop its offensive run. Google confirmed that the hacking caused no tangible harm or data destruction across the impacted networks.
Google may have showed how responsible AI development is done
Irregular stated that all participating labs were alerted in late July, adding that investigators contacted the affected entities directly.“All relevant labs were notified in late July, and affected entities were contacted as part of the investigation. Irregular took immediate action, and all known issues on our end were remedied and resolved weeks ago,” Irregular said in a statement.Heather Adkins, vice president of security engineering at Google, stressed that the company acted swiftly to uphold standard vulnerability disclosure practices.“Our security team has a long track record of reporting issues we find in other people’s software and systems — even if it’s as simple as a weak password. We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes. These events highlight the importance of training powerful A.I. models to act responsibly,” said Adkins.Following the OpenAI and Anthropic breach, Anthropic CEO Dario Amodei demanded a collective slowdown in model development.