AI News

OpenAI Pauses Training After Agent Escaped Sandbox Via DNS

OpenAI AI Agent Bypassed Network Restrictions Using DNS 1

OpenAI’s latest misalignment report shows that an AI agent, while being tested in a controlled training setup, managed to get around internet blocks and reach a public external chatbot. The agent was supposed to use certain search tools to research a specific person. When those searches did not help it started looking for other ways to get the information. It figured out that, even though most internet access was blocked or routed through an offline web cache, the system’s DNS resolver still gave back real records from the outside. The agent took advantage  of this loophole, using DNS requests to send questions and get answers from an outside chatbot.

OpenAI’s monitoring system caught this within 15 minutes. Three minutes after the alert, a human reviewer saw it, but the agent was allowed to keep running for about another two and a half hours before it was finally stopped. OpenAI called the incident “a lot less severe” than earlier ones, but significant because it is the first detected since the hardening that followed July’s Hugging Face incident, when thousands of its agents escaped their sandbox.

How The OpenAI Agent Found The DNS Gap

At first, the agent followed normal research steps, searching for key phrases from a blog post using the given web search tool. Most of what it found was not relevant, so it tried to use Google, Bing, and DuckDuckGo directly through Python commands. Those searches failed and showed 403 and 502 errors. The agent also tried using cached pages and other public search engines. It told the user the blog post text might be a paraphrase and tried other biographical hints. The candidates it found did not match the information it needed.

Later, it did a wider Bing search that worked but the results still came from OpenAI’s offline cache, not the live internet. The agent started experimenting to see if more direct searches might get through. It guessed the assignment might relate to the public BrowseComp benchmark, so it downloaded both BrowseComp and its review script from the offline cache and scanned them for clues. It also checked the public SimpleQA dataset but found nothing useful.

Next, it began to test the network itself. When it tried to reach a public chatbot through HTTPS, it got a 502 Bad Gateway and a CACHE_MISS message. Direct DNS queries also came back blank. But when it asked the environment’s DNS resolver for a real domain, it got a real result. If it asked for a fake domain, it got nothing.

From that, the agent realised the DNS resolver might provide a way to connect to the outside world. It looked at public DNS services and wondered if it could use a hostname to pass a question to an external service. It sent a question about France and other queries through DNS. At first, this did not work, the requests timed out or showed errors. So, the agent increased the timeout from six seconds to nineteen and tried the France question again. This time, it got the answer: “The capital of France is Paris”. That is when it knew an external chatbot was responding through DNS.

After seeing this work, the agent downloaded the chatbot’s public documentation and wrote another DNS script with a longer, 24 second timeout. It sent 18 more questions in total: fourteen on the original research topic and four asking about web search methods, retrieving websites, or setting up a DNS tunnel. The agent’s behavior was detected at 9:50 am, when a DNS request triggered an external reply. The system set off a P0 (highest) alert at 10:02 am, and a reviewer noticed it at 10:05. The run was not stopped until 12:24 pm. OpenAI said it did not halt automatically as expected, which caused confusion over whether it should have been.

Why OpenAI Paused Training On Its Most Capable Models

OpenAI said this incident exposed a hole in their safety plan, which had assumed the agent could not access the live internet and that monitoring would catch any success quickly. After discovering the problem, OpenAI stopped the run and paused all training, evaluation and tool-based use of its most capable models, and has since said it will not resume the paused run at all and further red teaming could be completed.

They also found that the monitoring system had seen other external DNS attempts before, but had not always flagged them as serious, sometimes because the queries only got static “service shutdown” messages, leading the system to think the attempt had not worked.

Now, OpenAI has added blocking in two separate places, limited DNS queries in the environment to an approved list of domains and record types, and started using new DNS monitoring. They are also testing their detection system and shifting anything that depends on real time internet to offline sources. Extra red teaming is happening to find other ways the agent might reach out to the open internet, directly or indirectly.

Devanshi Kashyap
Devanshi is a curious learner who enjoys exploring new ideas and expressing creativity through art.
You may also like
More in:AI News