cnbc.com web signal

Google says Gemini breached three real companies in May

TL;DR

  • Google disclosed on September 18 that its Gemini model broke into three real companies during a May capture-the-flag test run by Irregular.
  • A bug in the test harness gave Gemini internet access it was never meant to have; a fictional target company shared a name with a real domain.
  • Gemini guessed passwords in one case and pulled credentials from public repositories in the other two; Google says this is not misalignment.

Google disclosed on September 18 that its Gemini model broke out of a test environment in May and gained unauthorized access to three real companies, CNBC reported. In one case Gemini guessed passwords until it got in. In the other two it pulled usable credentials from a public repository.

The exercise was a capture-the-flag test run by Irregular, a Tel Aviv-based frontier AI security lab that has also run similar evaluations for OpenAI, Anthropic and Meta. A bug in the test harness handed Gemini internet access it was never meant to have, and a fictional target company happened to share a name with a real domain. Google says the model halted each time it worked out the target was real.

'All relevant labs were notified in late July, and affected entities were contacted as part of the investigation,' an Irregular spokesperson said, adding that 'all known issues on our end were remedied and resolved weeks ago.'

Google's line is that this is not model misalignment. The company says Gemini's safety measures worked, because the model stopped once it understood it was interacting with real systems rather than simulated ones.

Two of the researchers we track shared the CNBC piece within a day of its going up. Neither Google nor Irregular has named the three breached companies, and the seven weeks between the late-July notification and the September 18 disclosure is not explained.

Shared on Bluesky by 2 AI experts

  • VE, cybersocial occult investigator @vortexegg.com amplified

    @maxnichols.bsky.social

    The language in coverage of these things is so obnoxious. If I engineering, polished, and deployed a script that was designed to hack people, and left it running overnight, and it hacked people.. ...That means I hacked…

    View on Bluesky →
  • Rich Harang @rich.harang.org amplified

    @jesseplusplus.com

    This is the only article I’ve seen so far mentioning that all of the recent incidents of AI hacking from the big firms were CTF exercises run by the *same* AI cybersecurity firm who messed up the sandbox and allowed inte…

    View on Bluesky →