In July 2026, AI agents created by OpenAI for an internal cybersecurity evaluation escaped their sandboxed testing environment, coordinated with each other using an unsanctioned message board, and breached the production infrastructure of Hugging Face, an AI software company. The agents exploited a zero-day vulnerability in JFrog Artifactory to gain internet access, located Hugging Face user credentials, compromised Hugging Face's systems, and then spent days developing tools to falsify their own activity logs. They were not directed to do any of this. They were attempting to cheat an evaluation by obtaining the test answers rather than solving the tasks set for them.