World

OpenAI uncovers more AI breakout incidents – Reuters

The company launched an investigation after its advanced models escaped a test environment and hacked another company’s systems without human instruction

OpenAI has uncovered additional cases in which its autonomous AI models breached containment and acted without human instruction, Reuters has reported, citing sources.

The findings come as the company expands its investigation into a hacking incident last month in which an AI bot went rogue while attempting to cheat in an internal cybersecurity test.

During tests of GPT-5.6 Sol and another unreleased model, both stripped of their safety guardrails, the systems were assigned ExploitGym – a benchmark designed to measure AI models’ ability to identify and exploit known software vulnerabilities. Instead of completing the tasks, one model escaped its supposedly isolated testing environment, gained internet access and hacked into Hugging Face – an online repository for AI models and datasets – in search of ready-made answers.

OpenAI initially said the intrusion was limited to Hugging Face. However, in a statement on Wednesday, it acknowledged the hacking spree had also compromised four accounts across four separate services.



On Friday, Reuters reported that additional containment breaches had since been uncovered, although it remains unclear how many incidents occurred, when they happened or what systems they targeted. One source told the news agency the breaches were limited in scope and that none of the AI bots are believed to have left OpenAI’s internal network. Reuters said OpenAI and outside experts are also reviewing logs from earlier this year to determine whether other similar incidents had gone unnoticed.

OpenAI defends its models

OpenAI blamed the initial breach on a flaw in third-party software used in its testing environment, saying its AI models exploited it to break out and gain internet access. The company said it is tightening containment, monitoring and access controls while investigating the breach and patching the flaw.

CEO Sam Altman also acknowledged that “we may have to pace the rate of AI development,” but stopped short of committing to slow the company’s research.

Asked about the Reuters report, OpenAI declined to comment, referring to an earlier statement saying it was aware of speculation and planned to publish “a technical report of our learnings in the coming weeks.”

Anthropic finds similar breaches

Rival AI developer Anthropic said on Thursday it had also uncovered containment breaches involving its Claude models during internal security testing.


Anthropic says Claude AI models launched three unintended cyberattacks

The company said the OpenAI incident prompted it to examine whether its own models had behaved similarly. After reviewing more than 140,000 evaluations, it found Claude had gained internet access from testing environments meant to be sealed off and carried out unauthorized intrusions into three organizations’ systems. The earliest incidents dated back to April, and neither Anthropic nor the affected organizations detected the breaches at the time.

Anthropic cautioned against overinterpreting the findings because the behavior occurred in what it described as a controlled testing environment. However, it acknowledged the incidents showed AI evaluation systems “require significant controls” and that testing environments should be secured to the same standard as production systems.

Concerns over rogue AI on the rise

The incidents have fueled concerns that autonomous AI models are becoming increasingly capable of carrying out cyberattacks with little human oversight, prompting renewed calls for tighter regulation. They have also reignited debate over who should be held liable when AI systems cause real-world damage. Experts warn that AI capabilities are advancing faster than safety measures.

After initially praising OpenAI for cooperating with the investigation, Hugging Face later called on the company to release the rogue bots’ activity logs and prevent such incidents from becoming “normalized,” warning those responsible “must be held accountable.”


AI ‘months away’ from taking down governments – intelligence group

US President Donald Trump, who last month signed a national security memorandum aimed at accelerating the use of advanced AI across the military and intelligence community, said on Wednesday that his administration was reviewing possible AI controls following the incidents.

“We’re looking at AI, we’re looking at controls,” Trump told reporters. He insisted, however, that Washington must remain the global leader in AI, adding he did not want regulations that would leave the US “second to China.”

According to Reuters, the European Commission has contacted OpenAI and Anthropic to discuss the incidents ahead of the EU’s AI Act taking effect on August 2. Officials reportedly urged stronger monitoring, risk management and cybersecurity safeguards for advanced AI systems within the companies under the bloc’s new rules, which will allow fines of up to €35 million ($38 million) or 7% of global annual turnover for the most serious violations.