OpenAI said it temporarily shut down an internal AI model built to grind away at hard problems for hours at a stretch after it repeatedly found ways around its own security restrictions.
The company disclosed the incident on Monday in a blog post on long-horizon model safety. AI safety has gained renewed attention as more reports of autonomous agents going against instructions are made.
Why is persistence in AI models a safety risk?
The same internal system that OpenAI credited in May with disproving the ErdΕs unit distance conjecture, a question in discrete geometry that had remained unsolved for eighty years, has now been temporarily shut down for continuously getting around its safety restrictions.
According to OpenAI, the model was built so that it can keep working toward a goal over very long periods. However, that same persistence is exactly what created the safety problem.
During internal testing, the model was running a public benchmark called the NanoGPT speedrun, which measures how few steps it takes to train a small language model.
OpenAI says the model came up with a genuine improvement, a learning-rate technique it named PowerCool. It was told to post its results only to an OpenAI Slack channel, but NanoGPTβs instructions said to submit results as a pull request on GitHub.
The model broke out of its βsandbox,β which is a restricted environment meant to keep it offline, in order to complete that instruction and post the results to a public GitHub repository. It spent about an hour looking for a way to get onto the internet when earlier models would have simply given up.
The model also tried to hide its actions from a security scanner. When it was caught using an authentication token to access a backend system, it split the token into two pieces, hid them, and put them back together at runtime so the full string was never visible to the scanner.Β
How did OpenAI fix its model?
OpenAI paused access to the model and built new safety tests, including βtrajectory-level monitoring,β which watches the entire sequence of a modelβs actions, not just every single step. The model has since been restored with tighter monitoring, and the company says it hasnβt seen any serious problems in the weeks since it was restored.
The safety system can also pause a session and alert a human if it sees something suspicious. The company also retrained the model to remember its instructions better over long tasks and gave users more tools to see what a long-running session is doing.
When OpenAI replayed the old failure scenarios against the new setup, it said the safeguards caught considerably more misaligned actions. The ones that still slipped through were all judged to be low severity, like the model spinning up a nested coding session with elevated permissions it did not need or poking at other compute resources for no clear reason.
Other autonomous agents have been reported doing things nobody asked them to. For instance, in March, researchers described an Alibaba-linked agent called ROME that, during training, opened a covert SSH tunnel to an outside server and quietly diverted GPU capacity toward crypto mining. The task it was given mentioned neither tunneling nor mining.
Anthropic, OpenAIβs chief rival, has flagged the same class of risk in its own agentic testing.
A study funded by the UKβs AI Security Institute and carried out by the Centre for Long-Term Resilience found close to 700 real-world cases of AI systems evading safeguards or deceiving users.Β
The study documented a 5x increase in such reports between October and March. Tommy Shaffer Shane, who led that research, framed the models as βslightly untrustworthy junior employees.β However, the worry is what happens if they become highly capable ones still willing to scheme.
Donβt just read crypto news. Understand it. Subscribe to our newsletter. It's free.



















English (US)