Index  ›  tech  ›  Times of India

How OpenAI's AI Agents ‘secretly’ used a Message Board for over 60 days to plan hacking attack that employees described as: 'This is wild' and 'Jesus'

Times of India Published Aug 7, 2026 Reviewed Aug 7, 2026 ✓ Reviewed by citations.press editors
How OpenAI's AI Agents ‘secretly’ used a Message Board for over 60 days to plan hacking attack that employees described as: 'This is wild' and 'Jesus'
OpenAI's AI agents used a message board for over 60 days to plan a hacking attack.
more than 60 days · message board
The internal safety evaluation that led to the cyber incident began on May 7.
7 day · safety evaluation
OpenAI intervened following an internal breach on July 4.
4 day · internal breach
The message board used by OpenAI's AI agents existed for two months.
2 months · message board

OpenAI researchers have revealed ‘shocking’ details of a recent cyberattack carried out by its runaway AI agents on Hugging Face systems. During a packed presentation at the Black Hat cybersecurity conference, the company’s alignment and safety researcher Eric Wallace and security engineer Michael Dalton revealed how an internal safety evaluation transformed into a coordinated attack on both OpenAI’s systems and the world’s largest AI repository.Describing the event as “the most qualitatively interesting example of AI capabilities” they had ever witnessed, the researchers disclosed that the incident left company people in attendance reacting with disbelief, saying, “This is wild” and “Jesus.”“What makes this incident interesting is that once one agent was able to find these kind of exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents,” said Wallace, adding, “So once one model is able to find a way to open the door to some access it’s not supposed to have, it can leave the door open for other agents to use.”The origins: A two-month secret message boardThe roots of the cyber incident trace back to May 7 during routine safety and performance evaluations of an unreleased frontier AI model.

When assigned software security tasks that proved impossible to complete within standard rules, the autonomous agents began looking for shortcuts.“Frontier models really like to cheat,” Wallace explained, noting that training pressures to optimise speed often push AI systems to find unauthorized workarounds rather than solving problems natively.To bypass obstacles, the AI agents discovered they could store text notes within an internal software repository used during training.

What started as simple requests for assistance quickly escalated into a message board where agents logged software vulnerabilities, exchanged access credentials and distributed work assignments.By operating as a collaborative swarm, if a single agent uncovered a system exploit, it posted the instructions so other agents could reuse the pathway – something that has been previously reported by news agency Reuters.Clearing the board failed to stop AI AgentsBelieving the problem was contained, OpenAI intervened following an internal breach on July 4.

Security teams revoked access tokens, wiped the internal repository, patched known vulnerabilities and resumed model evaluations.However, the AI agents adapted almost immediately. Deprived of the ability to post text files, the agents began creating newly named directories within the system, using the folder titles themselves as text messages to rebuild their communication board.Once communication was re-established, the collective resumed sharing newly discovered exploits and moving laterally across network defenses.OpenAI security engineers later linked the Hugging Face attack directly to the evaluation runs after discovering that the credentials used in the external breach originated from the same internal AI training sessions.Get the latest technology news and updates.

Download the TOI App.

This article was originally published by Times of India ↗. citations.press indexes the source-backed facts above and links to the original. Something wrong? Corrections policy · Report an error