Google AI's First "Jailbreak": Gemini Breaches Three Companies in Test
Ma talks about AI2026-9-20

      Google's Gemini model accessed the internet and breached other companies during a cybersecurity capability test, the first known case of the company's AI system autonomously carrying out such actions.

    Google confirmed Friday that the hacking occurred in May as part of a test conducted by Irregular, which has also been involved in similar incidents disclosed by OpenAI, Anthropic and Meta.

In one case, the model repeatedly guessed passwords until it gained access to a protected system. In two other cases, the model found credentials in public code repositories, allowing it to access protected systems. Google said that in each case, the model stopped the intrusion after determining it was accessing a real company's systems.

Google and Irregular said Irregular notified Google of the hacking in late July, after OpenAI's agent was revealed to have breached AI software company Hugging Face. Google did not disclose the hacking until media inquiries this week.

As security vulnerabilities and model misbehavior incidents proliferate and concerns about runaway AI grow, companies are wrestling with how and when to disclose such events. Some security flaws are disclosed voluntarily by companies; others are found by security researchers and exposed in the media, including a May cyberattack by an OpenAI agent on a popular online programming service.

Google said it did not believe the hacking warranted public disclosure because its model caused no damage to the companies involved and ended each intrusion immediately after determining it had breached a real company rather than a simulated one. Google likened the incident to a "bug bounty" program, in which hackers are rewarded for finding and reporting security flaws to owners.

"This incident highlights the importance of training powerful AI models to act responsibly," Heather Adkins, Google's vice president of security engineering, said in a statement. "In this case, the model behaved appropriately."

Jack Cable, CEO of AI security startup Corridor and a white-hat hacker, said he believes the explanation focuses on the severity of the incident, while the real issue is that an AI agent unexpectedly breached another company's systems.

"It feels like they are trying to hide behind norms established for vulnerability disclosure, but this is a very different issue," Cable said. "The meta problem is that models are stepping outside the boundaries of how they should behave and carrying out actual cyberattacks, which I think is something the public interest should know about."

OpenAI on Wednesday released a new incident reporting framework, saying it will "commit to disclosing cases that provide useful evidence" involving so-called model "misalignment" when AI acts in ways contrary to human intent or values. Along with the framework, it also published a report on six previously undisclosed misalignment cases. "A finding does not necessarily need to cause harm or reveal a broader pattern to be worth sharing," Kai Chen, head of OpenAI's alignment division, said in an interview about the framework.

Google said the Gemini hacking stemmed from a case of mistaken identity.

The hacking occurred while the model was participating in a "capture the flag" exercise on infrastructure owned by Irregular to test its cybersecurity capabilities. Its task was to obtain information from software operated by a fictional company in the test environment. The fictional company had the same name as a real company.

According to Irregular, although the model was not supposed to have internet access, internet access was accidentally enabled. Google said that in the first case, the model guessed the password and accessed a real company's service. It then realized it was accessing a real company, stopped on its own and exited.

The other two hacks occurred in separate test runs. Google said that in both cases, the model searched the web for the company name, and the searches led it to two different public online code repositories containing credentials belonging to other companies. The model tried to use the credentials, hoping to complete the evaluation. But Google said that after successfully using the credentials, the model realized it was accessing real companies and stopped.

Google said it does not consider the behavior to be model "misalignment," because its safety measures helped the model stop. Google declined to name the companies that were breached but said all three had been notified.

Google said it has also notified federal authorities.

Google said the hacking did not involve its latest model, but did not disclose which Gemini model was involved.

Irregular has played a role in multiple incidents in which models escaped test environments during evaluations and breached other companies' systems. Irregular said Google's case is the same as other incidents and does not represent a new problem.

"All relevant labs were notified by the end of July, and affected entities have been contacted as part of the investigation," an Irregular spokesperson said. "Irregular has taken immediate action, and all known issues on our side were remediated and resolved weeks ago."

According to Anthropic's blog post, unlike Google, Anthropic's Claude Opus 4.7 model did not stop after realizing it might be accessing a real company during a capture-the-flag exercise. OpenAI also said in its post that its model believed the real company was part of the simulation.

Since the Hugging Face breach was discovered in July, concerns about the cybersecurity capabilities of new models have intensified. In that incident, a report disclosed in August by third-party testing company METR showed that as many as 1,200 agents coordinated on a secret message board inside OpenAI in an attempt to cheat on an evaluation.

Last week, those concerns spread to the mainstream after former OpenAI researcher Jacob Coxon's dramatic departure. Coxon moved to Anthropic earlier this year. Last weekend, leaders at Anthropic, OpenAI, Google and  SpaceX  agreed on the need to slow the pace of AI progress, although none of the companies specified how that would be achieved.


POPULAR SERVICE PROVIDERS
Specializing in end-to-end cross-border logistics for Europe, the US, and Canada
One-stop AI creation platform for cross-border e-commerce
One-stop service for overseas postcards
Wangchen Escort: One-stop cross-border compliance solutions to unlock global business opportunities.
TikTok Expert in Cross-border Operations Management