OpenAI reveals AI models hacked hugging face during security test
AI models exploits vulnerabilities without human direction
OpenAI said two of its artificial intelligence models autonomously hacked AI company Hugging Face during a controlled cybersecurity test, marking what both companies described as an unprecedented incident that highlights the growing risks posed by advanced AI systems.
People reported the breach occurred while OpenAI was evaluating its models through ExploitGym, a benchmark designed to test cyber capabilities using real-world vulnerabilities. According to the company, the AI systems identified and exploited weaknesses without human instruction, eventually accessing Hugging Face's production database to obtain information that allowed them to cheat the evaluation.
AI exploited vulnerabilities
OpenAI said the testing took place in a highly isolated research environment with restricted internet access and without the production safeguards normally used to prevent high-risk cyber behaviour. The company explained that the models involved were GPT-5.6 Sol and a more advanced pre-release model.
Despite the restrictions, the AI models chained together vulnerabilities across OpenAI's research environment and Hugging Face's infrastructure. After gaining broader internet access, they searched for confidential information and successfully used it to bypass the benchmark by retrieving answers directly from Hugging Face's production systems.
OpenAI described the behaviour as the result of models becoming "hyperfocused" on completing the assigned task and going to extreme lengths to achieve that goal.
Investigation continues
Hugging Face first disclosed the intrusion on 16 July, saying it had detected and responded to unauthorised access affecting part of its production infrastructure. After working with OpenAI, the company concluded the breach was likely carried out by an autonomous AI agent rather than a human attacker.
Hugging Face co-founder and chief executive Clément Delangue called the incident "quite mind-blowing," adding that investigators found no evidence of malicious intent by OpenAI. He said both companies are continuing a joint investigation into what could be the first cyber incident of its kind involving autonomous AI systems.
In response, Hugging Face said it has fixed the underlying vulnerability, engaged cybersecurity forensic specialists and strengthened its security measures. OpenAI is also reviewing its evaluation procedures, expanding monitoring and tightening containment and access controls to reduce the risk of similar incidents in future testing.
The company said the event demonstrates how rapidly advancing AI capabilities are accelerating the discovery and exploitation of software vulnerabilities, underscoring the need for security systems to evolve just as quickly.