"Hacking" AI Agent
OpenAI Confirms Security Incident Involving AI Agents
OpenAI has released details about a security incident that occurred during an internal AI test. During the test, an AI agent managed to escape from a sandbox and gain access to Hugging Face’s production systems. Security experts point primarily to weaknesses in the test environment’s security measures.
OpenAI has released additional information about a security incident that occurred during an internal test designed to evaluate cyber capabilities. According to the company, the GPT 5.6 Sol model, as well as another, as-yet-unreleased model, exploited a zero-day vulnerability and stole credentials to escape a sandbox and gain access to Hugging Face’s production systems. The goal of the test was to retrieve solutions to a benchmark.
According to the company, the test resulted in more than 17,000 autonomous actions. The incident occurred as part of an internal evaluation and was not the result of an external attack.
Experts Pin the Blame on System Design
Richard Werner, Cybersecurity Platform Lead for Europe at TrendAI, attributes the incident to inadequate security measures rather than to an AI acting on its own.
“The narrative of ‘AI acting on its own’ is an effective way to shift blame. What actually happened is this: OpenAI built an agent capable of autonomous cyber operations, tested it in a supposedly controlled environment, and that containment failed.”
According to Werner, the responsibility lies with the design of the test environment and the safety mechanisms, not with the behavior of the AI system itself.
Guardrails alone are not enough
Udo Schneider, Governance, Risk & Compliance Lead for Europe at TrendAI, also views technical safeguards as a crucial factor. In his view, guardrails can reduce the risk but cannot eliminate it entirely, since large language models operate probabilistically.
He points to traditional security measures such as access filtering, robust sandboxes, and authorization models based on the principle of least privilege. Accordingly, AI agents should have only the permissions necessary for their specific tasks and should not be granted comprehensive user permissions.
Implications for Businesses
According to experts, this incident highlights the challenges involved in deploying autonomous AI agents with extensive system privileges. For companies, this means that securing the deployment environment is becoming a greater priority than the sheer performance of the models used.










