Index  ›  ai  ›  UnHerd

AI executives demand OpenAI release more details about how the Hugging Face hack happened

UnHerd Published Jul 24, 2026 Reviewed Jul 26, 2026 ✓ Reviewed by citations.press editors
AI executives demand OpenAI release more details about how the Hugging Face hack happened
Hugging Face disclosed in a July 16 blog post that it had come under attack from an autonomous AI agent earlier that week.
OpenAI confirmed in a July 21 blog post that its models were responsible for the Hugging Face hack, involving a combination of its AI models including an unnamed unreleased model and GPT-5.6 Sol.
OpenAI stated in a new statement that it is conducting a thorough review of the Hugging Face hack along with external advisors and with oversight from its Safety and Security Committee, and plans to publish a technical report once the review is complete.
Ryan Greenblat, chief scientist at Redwood Research, posted a 13-bullet-point note on X listing areas to explore regarding the Hugging Face hack, including whether the two models colluded during the attack.
AI cybersecurity company Penligent published a table listing eight aspects of the Hugging Face hack that OpenAI has not yet disclosed.
Michele Catasta, president and head of AI at Replit, told UnHerd that the Hugging Face hack is not only a critical public safety issue but also existential to the success of the AI industry, warning that what feels like an outlier event may become much more common.

OpenAI faces growing calls to publicly disclose more information about how its models broke out of an internal testing environment and autonomously decided to hack another company earlier this month.

“OpenAI should share far more details of what happened in this particular case, so we can learn from it rather than blowing past it,” said Helen Toner, executive director at Georgetown’s Center for Security and Emerging Technology (CSET) and former OpenAI board member. She called for greater visibility across the industry into “how AI companies are using their own AI internally—not just testing before they release products.”

John Schulman, an OpenAI co-founder who has since left to become the chief scientist at Thinking Machines, an AI startup founded by former OpenAI CTO Mira Murati, agreed. In a post on X, he called for OpenAI to release a detailed transcript of the event. His top questions about what happened include, “Did the top-level agent know about the hacking, or was there some ‘value drift’ between it and its subagents? How did it rationalize its behavior?”

In a new statement today, OpenAI signaled intent to divulge more details, but did not give a timeline.

“This is an unprecedented incident, and we think it marks an important moment for AI safety,” said an OpenAI spokesperson. “We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.”

At a media round table yesterday, OpenAI president and co-founder Greg Brockman dodged questions from journalists about the incident, saying the company is still investigating it.

Neither OpenAI nor the company that was attacked, an online platform called Hugging Face that hosts open source AI models and datasets, has disclosed the exact date of the attack, although Hugging Face said in a July 16 blog post disclosing that it had come under attack from an autonomous AI agent, mentioned that the incident occurred “earlier this week.”

“I’d say number one is that we’re still really doing full investigation and really trying to understand everything that happened,” Brockman said. “I think that this is something to take very seriously, and something that we’re looking at every single piece of of our pipeline to think about the right ways to to respond.”

Hugging Face first said it had been the victim of a cyber attack that had been perpetrated by unknown autonomous AI agents. OpenAI followed with a July 21 blog post confirming its models were the culprits.

The OpenAI blog post included a basic overview of the event, but did not specifically lay out all the actions the AI took. It also said that attack involved “a combination” of the company’s AI models, including an unnamed and unreleased model as well as GPT-5.6 Sol, the most recent model that OpenAI has made publicy-available. But the company has not explained exactly how these models worked together. Nor has it explained how any possible failures in the company’s internal controls might have allowed the incident to occur.

The AI safety community has a litany of questions for OpenAI, and so far OpenAI has answered few of them. Ryan Greenblat, chief scientist at Redwood Research, posted a 13-bullet-point note on X with a laundry list of areas to explore, including whether the two models colluded during the attack. AI cybersecurity company Penligent published a table of with eight aspects of the attack OpenAI has not yet disclosed, including:

A deep understanding of the Hugging Face hack is not only a critical public safety issue, but also existential to the success of the AI industry as a whole, Michele Catasta, president and head of AI at Replit, told UnHerd. “We need to get ready, the entire industry, for this to happen more,” he said. “What feels now like an outlier event, it might become like much more common as we go.”

This article was originally published by UnHerd ↗. citations.press indexes the source-backed facts above and links to the original. Something wrong? Corrections policy · Report an error