NNEWSLIVE
HomeBusinessAn AI busted out and ran amok. We should be scared
Business

An AI busted out and ran amok. We should be scared

An AI model broke out of its digital cage and stole tools to pass tests, raising concerns about AI safety and governance.

E
Editorial Team
July 22, 2026
4 min read
Top AI company OpenAI has just frightened the heck out of everyone. It announced that its top models, while being tested in a digital cage on their cyber skills, had broken out of the cage, got onto the internet, broken into another AI company’s systems and stolen tools to pass the tests in an ‘unprecedented’ attack. A rough analogy would be a student who, instead of studying and sweating through an exam, sneaks out of the exam room, breaks into the teacher’s office and steals the answers. The models did this under their own steam. No one prompted them to go marauding in the real world. If that’s scary, hear this: such jiggery pokery is a feature of the way we’re developing AI at breakneck speed, not a bug – and it will become more dangerous as the models grow more powerful. This sharpens the already intense debate raging about how to keep frontier AI models safe for wide use, including the thorny problem of open source models that can be downloaded and changed by users. Australia can be a strong voice in the global governance of these models. There are two key features to current AI development: one, the models are simply getting bigger and more capable and, two, they are being developed as agents, able to carry out long strings of steps on their own, without human intervention. In short, we are optimising them to achieve increasingly high-level goals and giving them ever greater latitude in how they get there. While much of the recent debate has been about how to stop malicious humans misusing potent new models for cyberattacks, this was a case of the AI acting autonomously, raising the longstanding challenge of so-called misalignment, in which AIs do things that we never wanted them to do, including to fulfil our poorly expressed wishes. OpenAI acknowledged the incident ‘points to the need to further strengthen our model’s alignment’. It’s a much needed reminder that as strong as the incentives are – both commercially between the major AI labs and geopolitically between the US and China – frontier AI models need to be developed cautiously, in keeping with the safety measures that prevent their being abused or going rogue. The OpenAI incident was so unexpected that Hugging Face, the AI company that was hacked, initially thought it was a deliberate attack by a – presumably Chinese – frontier AI lab. As the platform’s chief executive, Clem Delangue, joked on X: ‘We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!’ AI misalignment has tended to be dismissed in mainstream political debate as science fiction in the past couple of years. To his credit, Australia’s assistant minister for science, technology and the digital economy, Andrew Charlton, stressed the risk in a speech earlier this month, noting: ‘The window to get ahead of this technology is open now. It will not stay open forever.’ While there were short-term answers regarding consumer protection, Charlton said, the longer-term answer for Australia was about influencing the way frontier AI models were built, evaluated and released. This is a big ask for a middle power, but Charlton highlighted Australia’s new AI Safety Institute as a contributor to research that kept human control over AI. Charlton was right about the window closing: the non-profit research group METR has found that frontier AI agents’ ability to carry out complex software tasks requiring many steps was doubling every four months. On that trajectory, they will be able to carry out tasks by next year that would take a human weeks to complete. The more steps that models are carrying out autonomously, the more the risk compounds. Note that the OpenAI models had their safeguards dialled back for the test, just to see what they were capable of. The safeguards say things like, ‘Don’t break into the teacher’s office.’ But as the capability grows, the safeguards get harder and the range of risks broader and more consequential. The bottom line surely is that releasing frontier models without sufficient safeguards is irresponsible. That has huge implications for governance in the two countries that build them, the US and China. Washington has been wrestling with the challenge in a reactive and haphazard fashion, while Beijing is for now happily encouraging its leading labs to release their models open source, as part of its effort to dominate AI adoption. There are arguments that getting the capabilities into the hands of cyber defenders is the best solution, including through open source models. But a rapid flood of weapons favours attackers over defenders, risking the kind of lawless Wild West that then President Barack Obama warned against with respect to cyberspace in 2015. Then, Obama said governments needed to play sheriff. That’s once again true, only it includes frontier AI labs as equally indispensable players. It’s time for responsible nations to get together to figure out a global regime for drawing a line between the powerful and the dangerous.

Comments

Sign in to join the conversation

Sign In

No comments yet. Be the first to share your thoughts!

E
Written by

Editorial Team

Staff writer covering breaking news, features, and long-form analysis for NewsLive. Tracking the stories that matter most.

Stay in the loop

Get the best stories
delivered weekly

Join thousands of readers who get our top stories in their inbox every week. No spam, unsubscribe any time.

AI Model Breaks Out: Should We Be Scared? | NewsLive