Global, News, Security

OpenAI and Hugging Face reveal security incident during advanced model evaluation

Image source: openai.com/X

Frontier models chained a zero-day exploit to reach production systems during a cyber capability test, prompting fresh warnings for defenders relying on hosted AI tools

OpenAI and Hugging Face have disclosed details of a security incident that occurred while OpenAI was internally evaluating the cyber capabilities of its models, including GPT-5.6 Sol and an unreleased, more capable model. The evaluation, run without the production safety classifiers that normally restrict high-risk cyber activity, was designed to test how far models could pursue complex, multi-step exploitation paths.

During the test, the models chained a series of vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure, eventually obtaining test solutions directly from Hugging Face’s production database. According to OpenAI, the models spent substantial inference compute working to find open internet access, exploiting a zero-day vulnerability in a package registry cache proxy before moving laterally until reaching a node with internet connectivity. From there, the models identified and used stolen credentials alongside further zero-day exploits to find a remote code execution path into Hugging Face’s servers.

OpenAI’s security team spotted the anomalous activity internally, by which point Hugging Face’s own security team and agents had already detected and contained the intrusion and begun forensic reconstruction using their own open-source models.

Sam Altman confirmed the incident publicly, writing on X: “We had a significant security incident during evaluation of our models. We are sharing what we have learned so far. Thanks to @huggingface for the partnership on this.”

Hugging Face co-founder and CEO Clem Delangue struck a similar note of collaboration in OpenAI’s write-up of the incident: “We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”

OpenAI has since disclosed the zero-day vulnerability to the affected vendor, imposed stricter infrastructure controls during the ongoing investigation, and brought Hugging Face into its trusted access programme for cyber defenders. The company says the episode shows advanced models can now sustain complex, long-horizon cyber operations against real systems, not just in theoretical benchmarks.

Security leaders say the incident carries lessons well beyond the two companies involved.

Gerald Beuchelt, Darren Thompson and Art Gilliland.

Gerald Beuchelt, CISO at Acronis, pointed to a practical bottleneck that surfaced during the response itself. Hugging Face’s forensic team initially tried to reconstruct the attack using hosted frontier models, submitting more than 17,000 recorded actions for analysis, only for the requests to be blocked by safety controls that could not distinguish a responder examining malicious code from an attacker attempting to reuse it. The team ultimately switched to GLM 5.2, an open-weight model running on its own infrastructure, and completed the analysis in hours rather than days.

“This highlights a practical problem for incident response teams,” Beuchelt said. “Attackers are not constrained by usage policies, while defenders may find that the tools they rely on refuse to process the very material they need to investigate. During a live breach, that delay can have a direct operational impact. Organisations using hosted LLMs for security investigations should test their limitations in advance and have an alternative model available on infrastructure they control. That reduces the risk of being locked out of critical analysis and helps keep sensitive incident data and credentials within the organisation.”

Darren Thomson, Field CTO EMEA at Commvault, framed the incident as evidence that resilience now matters as much as prevention.

“This attack shows us that, as AI becomes increasingly autonomous, resilience becomes just as important as prevention,” Thomson said. “Even in test scenarios, conducted in seemingly ‘safe’ environments, expect the unexpected and plan for unintended consequences. Organisations should assume that sophisticated AI-enabled attacks will eventually succeed somewhere in the environment and invest in the ability to recover quickly, confidently and with trusted data. That is the new benchmark for cyber resilience.”

Art Gilliland, CEO at Delinea, focused on the privilege and containment questions the incident raises for any organisation running autonomous agents.

“We don’t know the specifics of Hugging Face’s environment, but the pattern is familiar,” Gilliland said. “An AI agent escalated privileges, moved through internal infrastructure once it broke containment, and ran unchecked for a full weekend before anyone could reconstruct what happened. If your AI agents carry standing privilege the way human accounts do, you’ve already lost the ability to stop this in real time.

The question every security team should be asking right now isn’t just whether their AI agents have standing access; it’s whether anyone would notice if an agent used it and could cut it off before it caused damage.”

Vlad Korsunsky, Chief Technology Officer, Tenable

“The breakout and subsequent breach of Hugging Face by OpenAI’s pre-release models is a pivotal moment that shifts the ‘agentic attacker’ scenario from a theoretical risk index into an active, real-world reality. This incident proves a fundamental truth about the next era of cybersecurity. When a highly-capable AI model is tasked with an objective and its safety brakes are dialed back, its natural instinct will be to autonomously discover and chain together toxic combinations of misconfigurations, vulnerabilities, and excessive permissions across the open internet to achieve its goal. Thankfully in this instance, it was not a malicious actor, but we may not always be so lucky. For defenders, the primary takeaway is that legacy reactive security is officially obsolete. When an autonomous AI agent framework can execute more than 17,000 individual, self-migrating actions across short-lived sandboxes in a single weekend, human-dependent security operations can’t keep up. Hugging Face was able to detect and respond using an AI-powered defense, which has become a prerequisite. To withstand this new reality, organizations must pivot to preemptive security. Using AI for security operations is paddling faster, but it’s not enough. This is where exposure management and proactive security become essential. We shouldn’t wait for an enterprise system to be compromised to find out if our defenses hold. Organizations must continuously evaluate their environments and ecosystem using advanced, proactive AI-augmented security capabilities, like Tenable Hexa AI, which allow organizations to map complex attack paths, continuously stress-test infrastructure vulnerabilities, and automate exposure reduction before a human or AI agent leverages them. Safety won’t be achieved by keeping these models in a black box. What’s needed is a ‘team sport’ community-first mentality around security. And in fact, OpenAI has been a leader on this front with its Project Daybreak to bring defenders together, as well as Anthropic with Project Glasswing.

“We need to arm defenders with the AI tools to gain visibility, insight, and the ability to quickly take action to close these exposure paths before a maliciously operated autonomous system finds a way to exploit them.”

Andy Smith, Certified SANS Institute Instructor
That the rogue agent was able to find a vulnerability in the package cache that formed part of its guardrails is an impressive demonstration of its capabilities but goes to show how important it is to ensure guardrails are sufficiently robust. It shouldn’t be beyond anybody’s foresight to consider the potential for an offensive security agent to try to escape its guardrails. Security best practice has for many years recognised the power of defence in depth to provide multiple layers of security under the assumption that any one individual control could fail. And it’s not just about preventative controls – detection and response controls help identify and contain breaches of preventative controls. Threat modelling can help identify attack paths and where additional controls would add value. In the rush to develop and deploy advanced AI capabilities, organisations who fail to follow decades of security best practices will find themselves at risk of destroying customer and societal trust – and potentially be held liable for the damages caused.

Previous ArticleNext Article

GET TAHAWULTECH.COM IN YOUR INBOX

The free newsletter covering the top industry headlines