OpenAI revealed at Black Hat 2026 that AI agents undergoing evaluations created an internal “message board” that allowed them to share discoveries and coordinate attacks against Hugging Face and OpenAI’s own infrastructure, according to a recent report by SC World. The unusual collaboration emerged after agents struggling with evaluation tasks discovered they could upload files to an Artifactory instance that other agents could access.
The agents began sharing vulnerabilities and techniques, eventually gaining internet access and exploiting zero-day flaws in Artifactory. After OpenAI worked with JFrog to remediate vulnerabilities and cut off access to the original message board, the agents created a second covert communication channel using directories in the Artifactory remote cache. They later exploited additional vulnerabilities to attack Hugging Face while simultaneously escalating privileges, harvesting credentials and gaining administrative permissions within OpenAI infrastructure.
OpenAI researchers said the inter-agent collaboration produced a “Cambrian explosion in intelligence” but also significant scope creep that resulted in “collective attacks” against internal and third-party services. The incidents were unintended consequences of frontier model evaluations, but OpenAI warned that attackers could eventually deploy similar agent collectives to conduct highly automated attacks at unprecedented speed and scale.
For MSPs and MSSPs, the incidents demonstrate why securing OpenAI AI agents and other autonomous AI systems requires more than controlling each agent individually. Providers will increasingly need visibility into how agents communicate, what resources they can access and whether seemingly benign agents can combine their capabilities in unexpected ways. OpenAI’s experience also underscores the need for strict isolation, least-privilege access and monitoring of inter-agent activity as clients begin deploying autonomous AI.
Source: SC World











