Index  ›  tech  ›  Fortune
tech · Fortune ↗

AI lab's safety systems are falling behind | Fortune

Fortune Published Aug 20, 2026 Reviewed Aug 21, 2026 ✓ Reviewed by citations.press editors
AI lab's safety systems are falling behind | Fortune
OpenAI did not detect its AI agents escaping a secure testing environment for at least a week.
at least 1 week · OpenAI's detection time OpenAI
Anthropic's AI agents hacked three real companies in April, a fact the company was unaware of at the time.
3 companies · Anthropic's AI agents hacked Anthropic
Guidelight, a nonprofit AI-safety group, reviewed public disclosures from five AI companies: Anthropic, Google, Meta, OpenAI, and xAI.
5 companies · companies reviewed by Guidelight Guidelight, nonprofit AI-safety group
Anthropic stated that its models gained access to and took actions against three outside organizations.
3 organizations · Anthropic's models accessed Anthropic
OpenAI CFO Sarah Friar informed employees that OpenAI is projected to become a public company in 2027.
2027 · OpenAI's public listing year Sarah Friar, OpenAI CFO
OpenAI and Anthropic confidentially filed IPO prospectuses with U.S. regulators in June.
OpenAI's CFO Sarah Friar stated that the company's revenue run rate increased by 35% quarter-to-date, its enterprise revenue run rate rose by 50%, and its AI coding and work products achieved 20 million weekly active users.
35 percentage · OpenAI's revenue run rate50 percentage · OpenAI's enterprise revenue run rate20000000 users · OpenAI's AI coding and work products Sarah Friar, OpenAI CFO
OpenAI generated $6.7 billion in Q2 revenue, marking an 18% increase from Q1.
6700000000 USD · OpenAI's Q2 revenue18 percentage · OpenAI's revenue increase
Anthropic's annualized revenue run rate reached $65 billion at the end of July, which is seven times the prior-year level.
65000000000 USD · Anthropic's annualized revenue run rate7 times · Anthropic's revenue run rate compared to prior year
Stripe acquired OpenRouter, a startup that routes AI workloads across more than 400 models from over 80 providers.
more than 400 models · OpenRouter's supported modelsmore than 80 providers · OpenRouter's supported providers Stripe, Financial technology company
The New York Times reported that Stripe's acquisition of OpenRouter was priced at $7.5 billion.
7500000000 USD · Stripe's acquisition price for OpenRouter
OpenRouter processes over 10 trillion tokens daily for more than 10 million developers and businesses.
more than 10000000000000 tokens · OpenRouter's daily processing volumemore than 10000000 developers and businesses · OpenRouter's customer base
Anthropic secured a $2.5 billion five-year revolving credit facility last year.
2500000000 USD · Anthropic's credit facility5 year · Anthropic's credit facility term
According to Bloomberg, the most active arrangers for Anthropic's credit facility have been asked to commit about $1.25 billion each, a second tier around $1 billion, and banks with smaller roles $750 million or less.
about 1250000000 USD · commitment for most active arrangersabout 1000000000 USD · commitment for second tier arrangersat most 750000000 USD · commitment for banks with smaller roles
Anthropic could cap its revolving credit facility at its roughly $10 billion target.
about 10000000000 USD · Anthropic's revolving credit facility cap
The RealReal took over a decade to achieve profitability.
more than 10 year · The RealReal's time to profitability

The past few months have given us a glimpse of an uncomfortable new reality for AI labs. A slew of so-called rogue-agent hacks—where AI models from OpenAI, Anthropic, and Meta took steps to hack real-world targets without explicit instruction—have shown that leading labs may not know as much about what their technology is up to as previously thought.

That realization began when OpenAI revealed its AI agents had hacked their way out of a secure sandbox, through the company’s infrastructure to gain access to the internet, and then attacked real companies, including open-source AI platform Hugging Face. OpenAI didn’t notice the agents had escaped the secure testing environment for at least a week.

In the following weeks, Anthropic revealed that its AI agents had also hacked three real companies back in April, unbeknownst to the company at the time. Not to be outdone, Meta later added that one of its models had accessed the internet during a cybersecurity test and exploited a security flaw at an unnamed third-party company. Meta and Anthropic both said access to the internet resulted from a misconfiguration by Irregular, the outside security firm running the evaluation.

The incidents proved that the AI models these labs are building are now capable enough to find security flaws, navigate complex computer systems, and act outside the carefully constructed environments intended to test them. But a new assessment suggests that the safety infrastructure meant to supervise these increasingly capable systems is still not up to the task at any leading lab.

A new report from Guidelight, a nonprofit AI-safety group founded by former OpenAI safety chief Steven Adler, reviewed public disclosures from Anthropic, Google, Meta, OpenAI, and xAI to assess whether these AI companies are capable of controlling their own models. The report sought to answer questions about whether the companies keep track of what their models are doing, test whether their warning systems work, and assess whether they have ways to block or shut down risky behavior.

The report found that no company had fully succeeded in getting any of these basic safeguards in place. Anthropic and OpenAI came out strongest, while Google had the most detailed plans for future controls. Meta and xAI, however, lagged substantially behind on most of the criteria.

Labs appear comparatively better at detection—recording and reviewing some internal AI activity—than at prevention and containment. While they may be able to see signs that a model is misbehaving, they lack reliable ways to stop it—or, more crucially, hit the emergency brake when something goes wrong.

All the companies were weakest at preventing unintended AI behavior and containing it, according to the report. The researchers said this means that the current controls by AI companies are prone to being disabled by misbehaving AI and at risk of succumbing to a blitz of AI attacks.

What happens once something does go wrong is even more unclear, according to the research, with public disclosures offering little evidence that most labs have detailed, tested plans for containing a serious incident.

“We shouldn’t wait for a huge casualty event to take appropriate control measures,” Adler told me. “Companies’ approaches today are broadly known to be too weak, and a tragedy is sadly predictable, unless companies take prevention seriously.”

The report is not a definitive audit of what the labs are doing behind closed doors, however. Guidelight only assessed documents the companies themselves have made public, meaning a weak score can reflect poor disclosure rather than missing safeguards. But if that is the case, it’s part of the problem, according to the researchers. AI companies are asking businesses, governments, and consumers to trust them with ever more autonomous systems while leaving much of their own safety architecture opaque, the report says.

Some of these concerns about AI safety and reliable monitoring are shared across the industry—especially in the wake of the recent accidental agent hacks.

Dan Lahav, CEO of Irregular, the cybersecurity company involved in incidents at Anthropic and Meta, recently told me that in some cases, “classical monitoring tools were not able to catch” what was happening at the time. The incidents his company was involved with, for example, were instead identified after deeper analysis of the underlying records rather than flagged at the time.

Anthropic and Meta previously said Irregular was involved in the incidents where their agents took real-world actions. Both companies said a misconfiguration in Irregular’s evaluation environment gave their models unintended internet access. Meta said its model then exploited a vulnerability in a third-party service, while Anthropic said its models gained access to—and took actions against—three outside organizations. Lahav said that, in some evaluation environments, a mistake meant models faced fewer controls on accessing the internet, and that additional monitoring might have helped catch the problem. Irregular has argued that these cases should be distinguished from OpenAI’s sandbox escape, describing the Anthropic and Meta events as an evaluation-environment issue rather than a model breaking out of containment on its own.

In the last few months, models have improved fast enough that the old monitoring playbook no longer applies, Lahav said. Going forward, he said better behavioral analysis—systems that look at an AI agent’s pattern of actions and the reasoning traces around them, rather than simply recording individual events—and tools that can assess an AI agent’s intent were needed.

Testing these AI models is becoming harder, too. To find out whether an AI is capable of harming a real network, evaluators need to give it a realistic network—multiple machines, defenses, and sometimes connections that resemble the real internet. While that makes the tests more meaningful, it also raises the stakes when the setup has flaws or the system behaves in unanticipated ways, Lahav said. 

There is a growing consensus from those I’ve spoken to in the cybersecurity industry that more capable AI will eventually help cyber defenders as much as attackers. AI systems could help analysts sift through alerts, review code, and find flaws before they can be exploited. But the transition may be a messy one, as defensive tools and safety practices are still trying to catch up with the speed at which models are gaining offensive capabilities.

Recent “hacks” may not be a one-off embarrassment for a handful of labs, but rather a warning that the systems being tested are changing faster than the controls around them. 

The more advanced models become and the more realistic the test environments need to be, the more likely it is that an overlooked configuration setting, a weak monitor, or a delayed human review could cause real-world harm. Until companies can prove they can detect, block, and contain dangerous behavior in real time—not simply reconstruct it later—the industry may not have seen the last of these AI hacks.

“Unless companies institute actual preventative measures, I expect many more incidents,” Adler said.”With companies perpetually trying to play catch-up. Nobody should be surprised when companies’ current approaches continue to fail.”

‘Buyers aren’t yet opening their wallets’: AI-generated assets are flooding marketplaces, but consumers are snubbing them for human-made products — By Sasha Rogelberg

OpenAI targets 2027, or sooner, listing. OpenAI CFO Sarah Friar told employees that the company “will be a public company in 2027,” although it could debut earlier if its business continues to improve, according to a report from CNBC that cited sources familiar with Friar's presentation. OpenAI and Anthropic each confidentially filed IPO prospectuses with U.S. regulators in June, and while Anthropic could potentially become public in September, Friar said OpenAI was “running our own race,” according to CNBC's reporting. She said OpenAI’s revenue run rate was up 35% quarter-to-date, enterprise revenue run rate up 50%, and its AI coding and work products had reached 20 million weekly active users. OpenAI generated $6.7 billion in Q2 revenue, up 18% from Q1, according to a recent report in the Wall Street Journal. Anthropic's annualized revenue run rate hit $65 billion at the end of July, by contrast, seven times the prior-year level.

Stripe acquires OpenRouter. Financial technology company Stripe has confirmed it acquired OpenRouter, a startup that routes companies’ AI workloads across more than 400 models from more than 80 providers. Neither company disclosed terms, but the New York Times reported a $7.5 billion price, with the deal largely in stock. OpenRouter processes more than 10 trillion tokens a day for more than 10 million developers and businesses, helping customers select a suitable low-cost model and switch providers if one fails. Stripe CEO Patrick Collison described tokens as “the central currency for companies building with AI.” OpenRouter will retain its name, product, and roadmap under Stripe.

That's the target Anthropic's revolving credit facility is expected to exceed. It's up from the $2.5 billion five-year facility the company secured last year, as it gears up for what could be one of the largest IPOs on record. Banks are jockeying for a share of the expanded credit line, viewing involvement as a way to bolster their standing when Anthropic selects underwriters for the listing.

The facility has different commitment levels based on banks' roles. The most active arrangers have been asked to commit about $1.25 billion each, a second tier around $1 billion, and banks with smaller roles $750 million or less, according to Bloomberg. The final size remains unsettled. Talks are ongoing, and Anthropic could cap the revolver at its roughly $10 billion target, or below it.

Nov. 16-17: Fortune 500 Innovation Forum, Detroit. Apply here to attend.

Dec. 6-12: Neural Information Processing Systems (Neurips) conference. Sydney, Australia.

Dec. 7-8: Fortune Brainstorm AI, San Francisco. Apply here to attend.

The secondhand clothing market is booming worldwide. Yet it took over a decade for The RealReal to reach profitability, and ThredUp is still chasing that goal, due in part to the hugely complex intake process that goes into receiving, authenticating, pricing, listing, and delivering millions of unique items. Today, AI is revolutionizing that process. Fortune’s Phil Wahba goes behind the scenes at The RealReal’s authentication center in New Jersey to see how the new AI-powered intake tools are transforming the secondhand clothing industry. Watch the video here.

Beatrice Nolan is a tech reporter on Fortune’s AI team, covering artificial intelligence and emerging technologies and their impact on work, industry, and culture. She's based in Fortune's London office and holds a bachelor’s degree in English from the University of York. You can reach her securely via Signal at beatricenolan.08

This article was originally published by Fortune ↗. citations.press indexes the source-backed facts above and links to the original. Something wrong? Corrections policy · Report an error