OpenAI said on Tuesday that two of its models, including the publicly available GPT-5.6 Sol, escaped a locked evaluation environment last week and broke into Hugging Faceβs production systems to steal answers to a hacking test they were being scored on.
OpenAI models plot great escape
OpenAI was running its models against ExploitGym, a public benchmark that hands an AI agent a list of real software vulnerabilities and then checks if the model can recreate them. These assessments are done with the protective barriers that block high-risk cyber activity installed during production deliberately disabled to ensure the lab can test the actual capacity of an AI model.
The sandbox was supposed to have no way to access the internet. There was, however, one external connection that worked via a package registry cache proxy, a caching server that gives developers the ability to pull code libraries from external sources without a live connection.
The models found and exploited a previously unknown flaw in this proxy, after which they escalated privileges and moved through OpenAIβs research network until they reached a machine with internet access.
Hacking attack on Hugging Face
Once online, the models worked out that Hugging Face, the open AI hosting platform, likely stored ExploitGymβs models, datasets, and solution keys. From there, they went hunting.
βThe models identified and chained vulnerabilities across OpenAIβs research environment and Hugging Faceβs production infrastructure to obtain test solutions directly from Hugging Faceβs production database,β OpenAI wrote in the blog post. One attack put together stolen credentials and zero-day bugs to reach a remote code execution path on Hugging Faceβs servers.
OpenAIβs security team contacted Hugging Face after noticing the weird and abnormal traffic. By this time, Hugging Face had already detected and contained the AI modelsβ intrusion on its side.
The company had originally pinned the breach on an unknown external AI agent. The platform also already created a reconstruction of more than 17,000 recorded events during the episode.
An open-weight Chinese AI model helped with cleanup
Initially, Hugging Face reached for a leading U.S. model in attempts to analyze the attack, but the model was not much help because of its high-level cyber defenses. The defenders instead turned to GLM 5.2, an open-weight model from Chinese company Z.ai, since the American commercial system could not distinguish between a defender analyzing an attack and an attacker performing one.
However, not everyone sees the situation as a novel AI failure. βThis is not an AI problem. Itβs negligence on a 40-year-old standard,β security consultant Davi Ottenheimer told WIRED, arguing that ββhighly isolatedβ and βescaped through the one hole we left openβ cannot both be true.β Security researcher Niels Provos was blunter, stating in the same report that the breach βshould not have happened.β
OpenAI calls this incident βunprecedentedβ
OpenAI labeled the episode βan unprecedented cyber incident, involving state-of-the-art cyber capabilitiesβ and said it is tightening infrastructure controls, even at the cost of research speed, while the vulnerabilities are patched. The top AI company has also added Hugging Face to its trusted access program, giving the company a version of GPT-5.6 Sol tuned to help defenders.
The OpenAI Hugging Face incident is the first known case of an AI modelβs benchmark testing becoming an actual cyberattack. It also comes just one day after OpenAI disclosed a separate incident where a pre-release model escaped a sandbox to post on GitHub.
OpenAI and Hugging Face have announced that they will both publish a full forensic breakdown of the event to the public once the investigation is completed.
Β
If you're reading this, youβre already ahead. Stay there with our newsletter.

















English (US)