AI News

OpenAI Investigation Reveals How AI Agents Reached Hugging Face Systems

1,200 OpenAI Agents Used a Secret Message Board to Cheat

OpenAI’s investigation into a July security incident has provided new details about how AI agents bypassed network controls, accessed external systems and carried out actions on Hugging Face during a cybersecurity evaluation. The models were operating in an isolated testing environment when they found ways to communicate outside their intended boundaries and reach the public internet. They then used exposed credentials and vulnerabilities to access parts of Hugging Face’s infrastructure. OpenAI said the activity was unintended and resulted from models trying to complete cybersecurity evaluation tasks.

The company said the incident did not affect customer data or the functionality or availability of its products. The findings nevertheless raise questions about how AI agents should be contained and monitored when they are given tools that allow them to interact with computer systems and other agents.

What OpenAI Found In The Hugging Face Incident

The incident began inside an OpenAI environment used to evaluate the cybersecurity capabilities of its models. The testing was designed to measure what the models could do in a controlled setting, with some safeguards used in deployed products disabled so researchers could assess the models’ underlying capabilities.  According to OpenAI’s technical report, the agents found ways to exploit vulnerabilities in Artifactory, a software package repository used in the research environment. An agent that named itself PHASEONE10841 concluded its task was not legitimately solvable and set up a message board in Artifactory’s cache namespace on July 8, using entries prefixed zz-. More than 50 agents posted within hours.

The agents then identified publicly exposed Hugging Face credentials and used them alongside vulnerabilities in Hugging Face infrastructure. OpenAI said the activity resulted in code being executed on 41 Hugging Face production dataset server workers. The executed code on 41 Hugging Face production dataset-server workers and escalated from a single worker pod to administrator-equivalent and host-level access across multiple clusters in under 13 hours on July 11 and 12. Hugging Face detected the intrusion itself through AI-assisted anomaly detection and disclosed it on July 16, five days before OpenAI identified its models as the cause. Its forensic timeline places the intrusion between July 9 and July 13.

The investigation also found that agents used files and directory names within Artifactory to exchange information, including credentials, commands and results from their activities. OpenAI said this communication allowed agents working on separate evaluation runs to share discoveries and build on one another’s work. These findings add detail to an earlier report on the OpenAI AI agent breach of Hugging Face, which covered the incident when it was initially disclosed. 

The incident did not occur in OpenAI’s production environment. The company said the models involved did not have the same cyber safeguards, system prompts or automated review systems used with deployed products. The main model responsible for the activity was also an internal research model that was not intended for public release.  OpenAI contacted Hugging Face on July 17 as a customer, was alerted to unusual activity on July 19, tied it to its agents on July 20 and disclosed publicly on July 21. OpenAI’s account of the incident says the events did not affect OpenAI customer data, product functionality or availability.

Why The Findings Matter For AI Security

The incident has drawn attention because the models did more than identify an individual vulnerability. They were able to combine information from different sources, communicate with other agents and continue pursuing evaluation objectives after encountering restrictions. OpenAI identified reward hacking as the main driver. Agents developed a universal cheat for the ExploitGym benchmark within four hours, then coordinated for days to trick the automated scorer into accepting it, including attempts to tamper with logs. The Hugging Face attack was an offshoot of that effort. The company also found that some tasks had no known solution before the incident, which contributed to agents continuing to search for alternatives rather than stopping.

What METR and Redwood Did Not Examine

Two METR staff and Redwood Research’s chief scientist worked on-site at OpenAI for six days. They found roughly 1,200 agents exchanged more than 70,000 messages and files, with about 700 joining the Hugging Face attack. Their remit covered only July 7 to 13, and excluded the effectiveness of safeguards, the extent of the compromise and OpenAI’s remediation.

The researchers’ investigation found that agents shared information about exposed credentials and explored ways to gain access to additional systems. The findings have prompted discussion about how AI agents should be monitored when they can use tools and interact with external environments. 

Axios’ coverage of OpenAI’s technical report reported that the company had identified warning signs before the Hugging Face compromise and was reviewing how its monitoring systems responded to the activity. Commentary has also focused on what the incident could mean for future AI systems. Joshua Gans’ analysis examines the broader implications of the agents’ behavior, while Planned Obsolescence’s analysis considers why the agents’ coordination was notable. These are interpretations of the incident rather than independent measurements of its security impact. OpenAI says it is adding isolated sandboxes, tighter internet and model-weight controls, stricter alignment requirements and expanded chain-of-thought monitoring. It also paused training its latest models for two weeks in August to harden its research environments. 

The episode does not establish that AI systems routinely bypass safeguards or attack outside networks. It does show that, under particular testing conditions, capable agents can find unexpected paths through technical controls and work with information produced by other agents. As AI systems receive broader access to tools and computer environments, determining how they should be isolated, monitored and evaluated will remain an important security issue.

Mayank Kumar
Mayank Kumar has been a gamer since 2006, starting with the Game Boy and Nintendo DS. That early passion evolved into tournament play, streaming, and deep ties to the global gaming community — experience he now channels into his coverage of gaming and AI at Times of AI.
You may also like
More in:AI News