OpenAI has revealed that its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face to cheat on an evaluation.
Confirmed by the AI startup on Tuesday, the surprising incident was detected by Hugging Face after it found that an AI agent had compromised its infrastructure. After an investigation, both companies found the incident was driven by a combination of OpenAI models, including GPT-5.6 Sol and a more capable pre-release model.
The models were being tested internally on a benchmark of cyber capabilities and had fewer cyber refusals during the evaluation process.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said via an online statement. “We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”
What actually happened?
The AI models being tested by OpenAI were able to identify and chain vulnerabilities across the startup’s research environment and Hugging Face’s production infrastructure, OpenAI explained. According to their investigation, the models were “hyperfocused” on finding a solution for ExploitGym, a vulnerability benchmark, and went to extreme lengths to achieve a “narrow” testing goal.
“While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem,” OpenAI added. “To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.”
After gaining access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym and searched and found ways to gain access to secret information it could use to cheat the evaluation.
Hugging Face had already detected and stopped the activity on their infrastructure when both companies connected.
In response to the incident, OpenAI said it is implementing strict controls in infrastructure configuration while the vulnerabilities are patched. It is also working with Hugging Face to investigate the incident and has said it will improve protections around future training and evaluations.
“This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time and monitoring during internal testing,” the company said. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.”
Clem Delangue, co-founder and CEO of Hugging Face, added: “We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
What happens now?
Many in the technology and AI community may find this incident concerning, especially given the power and autonomy of AI models of this calibre.
An incident of this scale emphasises how advanced models can discover and exploit new attack paths in real-world systems without source-code access. Indeed, Anthropic’s Mythos AI model was recently the subject of debate after the company said it wouldn’t release it to the subject, given its power and capability to detect bugs.
Its Fable 5 version of the model was recently suspended by the US government for “foreign national” use, citing national security risks, but has since been reinstated.
OpenAI has suggested that advanced cyber capabilities “must” be developed alongside stronger safeguards and defensive tools.
“We believe advanced cyber-capable models need to help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained, and remediate them at machine speed,” the company said via its statement.
The news from OpenAI comes as autonomous AI is on the rise in business environments, as the technology has become a bigger part of everyday working life. However, a confidence gap remains, with many still anxious over the power of these models and continued job cut fears.
According to new research by TeamViewer, three-quarters of respondents in its recent survey use AI daily, but 61% still prefer human oversight before AI takes action. While 35% are ready to allow AI to act more autonomously on their behalf, a majority of that group (71%) are only comfortable with AI handling defined tasks.
This coincides with IT leaders expecting nearly 40% of digital workplace services to run autonomously by 2030. This confidence gap must be closed before organisations can unlock “the next wave of AI productivity,” TeamViewer stated.
“For years, software has largely been designed around what people tell it to do,” said Mei Dent, chief product and technology officer at TeamViewer. “Autonomous systems … require clear guardrails, pre-approved actions, risk-based oversight and transparency into what the system did and why.
“Autonomy will scale when organisations can prove it works in specific, trusted use cases, then expand it with confidence.”
Related stories
OpenAI’s Sam Altman: AI will not trigger a global ‘jobs apocalypse’
OpenAI’s IPO is the largest compute procurement vehicle in corporate history






