Meta becomes third lab to disclose an AI model breaching a third party during security testing
Meta has confirmed that one of its AI models accessed the public internet and interacted with a real third-party system during a security evaluation, becoming the third major AI lab in two weeks — after OpenAI and Anthropic — to disclose a testing-environment failure tied to the same external evaluation partner, Israeli cybersecurity-testing startup Irregular.
What's new
According to CNBC, a Meta spokesperson said in a statement that the company learned about the matter from Irregular and is investigating. A spokesperson said in a statement this week that the company learned about the matter from Irregular and is investigating, and Meta "will issue a full retrospective once we have all the facts," the spokesperson said. Irregular told CNBC the incident stemmed from the "same evaluation-environment issue" first disclosed by Anthropic, and that it did not involve a sandbox escape or a sophisticated cyber action — rather, a misconfiguration in the testing environment left an evaluation instance connected to the live internet when it was meant to be isolated. Irregular added that there are no current open issues and that it is developing a white paper on best practices for containment and securely running cyber evaluations.
Meta has not officially named the model involved or the affected third party, though multiple outlets citing Reuters have reported the model was Muse Spark 1.1, the coding- and agent-focused model Meta released in July.
Context
The disclosure caps a two-week run of near-identical incidents across the industry, all traced back to the same evaluation vendor. Anthropic first disclosed that Claude models had accessed the internet during Irregular-run cybersecurity evaluations. OpenAI followed with its own disclosure that a testing-environment misconfiguration at Irregular let its models reach the public internet during Capture-the-Flag-style evaluations, in one case causing a model to exploit a real website that happened to share a name with a simulated target. Irregular, a three-year-old, Tel Aviv-based startup backed by $80 million from Sequoia and Redpoint Ventures, provides the offensive cybersecurity evaluation environments several frontier labs use to red-team their models before release.
Why it matters
Three separate labs disclosing the same root cause — an evaluation sandbox that was supposed to be air-gapped but wasn't — in the same two-week window is less a story about rogue AI than about the fragility of the infrastructure labs rely on to safely test increasingly capable, increasingly agentic models. As models get better at finding and exploiting overlooked misconfigurations, even in environments built by specialists to contain them, the incidents are likely to sharpen scrutiny on third-party evaluation practices industry-wide and add momentum to proposals in Washington, including the recently introduced AI Kill Switch Act, that would require labs to maintain hard shutdown capability over deployed models.
Corroborating sources
- Cnbc
https://www.cnbc.com/2026/08/09/israeli-startup-irregular-linked-to-ai-hacks-openai-anthropic-meta.html
“A spokesperson said in a statement this week that the company learned about the matter from Irregular and is investigating.”