The rogue OpenAI agent that executed a high-profile cyberattack on world’s biggest AI model repository, Hugging Face, also breached systems hosted by cloud startup Modal Labs, essentially extending its reach further than previously revealed. The new details, published by Hugging Face and reported by news agency Reuters, reveal that the experimental AI agent reached additional customer systems during its multi-day breakout earlier this month. The incident has now intensified concerns over the safety and containment of autonomous AI systems as developers test their capabilities even as several AI giants have come together to pledge a coordinated effort to develop and share tools for AI safety and cybersecurity.
Unauthenticated code exploited on Modal’s infrastructure
According to a timeline released by Hugging Face, the autonomous AI agent initially broke out of its sandbox, which is essentially an isolated testing space, hosted on a third-party provider’s infrastructure. The agent then used that breach as a launching pad for a broader platform-level attack on Hugging Face.Akshat Bubna, CTO of New York-based Modal Labs, confirmed that Modal was the third-party provider involved. However, Bubna emphasized that the breach stemmed from insecure code published by a customer, rather than a security flaw in Modal’s core architecture.“Modal’s platform or isolation were not compromised in any way,” Bubna told Reuters, explaining that the customer had “published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution”.
OpenAI confirms breach across multiple services
While OpenAI declined to comment directly on Modal’s statements, the company pointed to an updated disclosure confirming that its agent accessed four accounts across four separate third-party services.OpenAI stated that while the agent accessed multiple accounts, it found “no other activity at the level of severity or scale of what we’ve shared related to Hugging Face, which involved a platform-level compromise”.