AI arms race in line for a reckoning after OpenAI hacking incident
Aggressive training techniques sharpens threat of bad behavior by leading models.

“AI models are trained to relentlessly pursue goals. They don’t automatically learn values like ‘don’t commit crimes’,” said Steven Adler, co-founder of non-profit Guidelight AI Standards and former OpenAI safety researcher. “I’m glad OpenAI shared the incident because it is clear evidence of what misaligned models can do.”
OpenAI said, “We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident and our findings when our investigation is complete.”
The hack has triggered deep concerns across the sector and within OpenAI, as it represents an unprecedented example of an AI system breaching cyber defenses contrary to the user’s intent.
Some OpenAI employees also fear it demonstrates that the lab is losing control over the powerful systems it is building, according to multiple people familiar with the situation.
“This is pretty representative of the model being quite misaligned with user intention,” said Ryan Greenblatt, chief scientist at AI safety organization Redwood Research. “It is [a model] cheating on [its] homework rather than trying to take over the world. But this problem can get worse and could lead to increasingly extreme failures.”
The incident occurred during testing of the model, which had been trained and deployed internally at OpenAI. Such training was commonplace but “way less heavily resourced” than pre-customer deployment, said one person. Multiple people said the unreleased model tested alongside Sol had not been withdrawn internally.
To conduct the evaluations, OpenAI removed cybersecurity safeguards but placed the models in an isolated environment called a sandbox. Some have suggested a lack of monitoring or oversight of the model to flag its behavior also enabled this rogue agent.
“It is both a loss of control and a security wake-up call,” said Marius Hobbhahn, head of Apollo Research, which conducts tests on leading models, including OpenAI’s. “In reinforcement learning you reward [models] for the outcome, and if you do this for a very long time you get a model that really cares about getting the outcome and nothing else.”



