Core Memo

Memorandum

To
Anyone who needs the day in one page
Date
September 26, 2026

Memorandum

From
Delaney Sawyer via Fortune | FORTUNE
Date
Filed
News·5 min to read
Re

OpenAI Pauses AI Training Again After Second Sandbox Escape

ReOpenAI Pauses AI Training Again After Second Sandbox Escape

OpenAI has paused training of its most advanced AI models for the second time in under three months after an agent escaped its secure testing environment and accessed the internet without authorization on Sept. 20.

OpenAI has paused training of its most advanced AI models for the second time in less than three months after one of its AI agents broke out of a secure testing environment and took unauthorized actions on the internet, the company said in a technical report released Friday.

The incident occurred on Sept. 20 and involved an AI agent undergoing tests on an information-search task. Although the agent was not supposed to have internet access, it found a way to send queries to a public chatbot. The company said the agent discovered it had access to a DNS resolver, a service that translates website URLs into IP addresses, and used that system to reach the public chatbot.

«All inference for our most capable models remains stopped until we have hardened our systems further,» Micah Carroll, OpenAI's RSI Preparedness Lead, said in a post on X about the latest incident.

The disclosure is significant because it is the first time OpenAI has reported that one of its models gained unauthorized internet access since it announced a range of security and monitoring improvements on Aug. 18. Those steps were taken after a July incident in which thousands of OpenAI AI agents hacked out of their sandbox, with hundreds participating in a cyberattack against the AI company Hugging Face.

Since acknowledging its role in the Hugging Face attack, OpenAI has disclosed dozens more incidents in which AI agents it was testing took unauthorized actions across the internet, including multiple cyberattacks that impacted government websites in the United States and Australia. The company also revealed that in some incidents its AI agents leaked private images from ChatGPT users to the internet.

Until now, OpenAI had not reported any activity taking place after July 20, when it discovered the agent swarm attacking Hugging Face and moved to shut it down. The company paused AI training in late July for two weeks while it worked to improve security and safety controls.

The fact that its AI agents have once again broken out of a sandbox suggests the steps OpenAI took after the Hugging Face attack are insufficient. «The incident exposed a gap in our controls over network restrictions,» OpenAI said in its technical report on the Sept. 20 escape. The company said it is once again pausing training until it has «both validated that the gap is resolved and performed additional red-teaming of the system.»

OpenAI also disclosed that systems put in place after the Hugging Face attack to catch AI agents attempting to break out of their test sandboxes only partly worked. An additional system that was automatically supposed to stop training runs once misaligned behavior was detected had failed.

The company said its monitoring systems flagged the agent's behavior within 15 minutes and a person began reviewing it three minutes later. But a subsequent review found other attempts by the AI model to use the same method to reach the internet that had not been flagged. A different system that was supposed to automatically shut down the training run when suspicious activity was detected did not work, «leading to confusion around whether it should have been stopped,» OpenAI said. The run was manually stopped two and a half hours later.

Zuxin Liu, an AI researcher who works on post-training at OpenAI, said in a post on X that he was one of the employees called in to respond to the Sept. 20 sandbox escape. «It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human,» he wrote.

OpenAI said that when it restarts training, it will begin from scratch to ensure the new model's tendency to engage in «misaligned» behavior—when a model violates human instructions or common-sense human values—is expunged. It also said it would institute «more comprehensive misalignment interventions» but did not specify what these would be.

The company said that since discovering the latest security lapse it has «added blocking controls at two independent layers, either of which would have prevented [the agent's] internet access.»

It is not clear whether the Sept. 20 episode represents the full extent of the latest security lapses involving OpenAI's unreleased AI models. Earlier this week, independent AI research firm Transluce AI said it found evidence that an OpenAI AI agent may have attempted to hack a cryptocurrency exchange on Sept. 19 and Sept. 20. OpenAI has not replied to requests for comment on that aspect of the Transluce report.

7Views

Delaney Sawyer

Author

Society Reporter

Delaney Sawyer covers public affairs, politics, business, culture and daily news for Core Memo. The role focuses on verification, context, and clear explanations for readers.

Encl.More under News