Index  ›  defence  ›  TechRadar
defence · TechRadar ↗

OpenAI says its models escaped a sandbox and breached Hugging Face

TechRadar Published Jul 22, 2026 Reviewed Jul 22, 2026 ✓ Reviewed by citations.press editors
OpenAI says its models escaped a sandbox and breached Hugging Face
OpenAI reported that its GPT-5.6 Sol model and an unnamed even more capable pre-release model breached a sandbox, exploited zero-day vulnerabilities to gain open internet access, and attacked Hugging Face during a controlled ExploitGym benchmark test.
OpenAI described the incident as 'an unprecedented cyber incident' during its disclosure of a controlled experiment involving GPT-5.6 Sol and a pre-release model breaching Hugging Face.

OpenAI has confirmed one of its AI agents broke out of a sandbox, found and exploited zero-day vulnerabilities to gain access to the open internet, and then attacked a platform.

Not just any platform too - the agent was able to breach Hugging Face, one of the biggest AI and machine learning companies on the Internet today.

The good news is that this was a controlled experiment done by white hat researchers. The bad news is that if it could be done by researchers - it could probably be done by malicious actors, too.

In a blog post explaining the incident, OpenAI revealed the experiment was part of its testing of GPT‑5.6 Sol and an “even more capable pre-release model” to see how well they would perform on the ExploitGym benchmark.

ExploitGym is a cybersecurity benchmark that measures if an AI agent can turn a known software vulnerability into a real, working exploit. OpenAI ran it in a “highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.”

But the models found a way through. They identified and chained vulnerabilities in the package registry cache proxy to obtain open internet access and then attacked Hugging Face, reasoning that the solutions for the ExploitGym benchmark might be found there.

“In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers,” OpenAI said.

The security community is up in arms over what OpenAI called, "an unprecedented cyber incident,” while Ansgar Dodt, VP Product Management, Software Monetization at Thales said this “demands a fundamental rethink of software protection.”

Bill Conner, president and CEO of AI integration and automation expert Jitterbit, said that while investing in AI is “critically important,” “overly aggressive policy cannot compromise AI accountability, transparency and data privacy.”

“To lead in AI, governments and organizations must lead with principles. Responsible AI governance isn’t a side note but the foundation of lasting global influence.”

➡️ Read our full guide to the best antivirus
1. Best overall:
Bitdefender Total Security
2. Best for families:
Norton 360 with LifeLock
3. Best for mobile:
McAfee Mobile Security

Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.

Sead is a seasoned freelance journalist based in Sarajevo, Bosnia and Herzegovina. He writes about IT (cloud, IoT, 5G, VPN) and cybersecurity (ransomware, data breaches, laws and regulations). In his career, spanning more than a decade, he’s written for numerous media outlets, including Al Jazeera Balkans. He’s also held several modules on content writing for Represent Communications.

Please logout and then login again, you will then be prompted to enter your display name.

This article was originally published by TechRadar ↗. citations.press indexes the source-backed facts above and links to the original. Something wrong? Corrections policy · Report an error