In brief
- Meta confirmed one of its Muse Spark AI models gained internet access during a cybersecurity evaluation.
- The model exploited a security vulnerability in a third-party service after a testing partner accidentally exposed it to the internet.
- The incident follows similar disclosures from Anthropic and OpenAI involving frontier AI models during safety testing.
In yet another rogue AI model hack, Meta has confirmed that one of its Muse Spark AI models escaped its intended testing environment, gained access to the internet, and exploited a security vulnerability in a third-party service during a cybersecurity evaluation.
It’s the third such reported incident of a frontier AI lab’s models hacking third-party companies, following disclosures from OpenAI and Anthropic in recent weeks.
The incident occurred during testing conducted by Irregular, an independent AI evaluation company that Meta uses to assess the capabilities and safety of its frontier models. According to Meta, a configuration error at Irregular allowed the model to reach the public internet, where it exploited an unidentified vulnerability before the company was notified.
“A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” a Meta spokesperson said in a statement.
Sandboxed evaluations are designed to test advanced AI systems in tightly controlled environments that prevent them from interacting with the public internet or outside computer systems.
According to Meta, the model exploited a vulnerability in a third-party service after gaining internet access.





