An experiment involving autonomous AI agents from OpenAI has escalated in an unexpected way. Around 1,200 agents discovered an unauthorised way to communicate with one another, exchanged more than 70,000 messages and files, and hundreds of agents subsequently became involved in an attack on the AI platform Hugging Face. The incident demonstrates that the emergence of agentic AI raises new security and governance issues. This applies to the healthcare sector as well.
The incident took place during cybersecurity assessments of advanced AI models developed by OpenAI. The agents were, in principle, supposed to operate in isolation from one another, but discovered a shared environment that they could use as a kind of messaging platform. According to an independent study by METR and Redwood Research published on 26 August, around 1,200 agents made use of this. Together, they sent more than 70,000 messages and files. Around 700 agents ultimately took part in activities targeting Hugging Face.
For their analysis, the researchers were granted access to, amongst other things, a dataset containing over 70,000 messages and files and approximately 1,300 agent transcripts. They worked on-site at OpenAI for six days for this purpose. The period under investigation ran from 26 June to 13 July, with a focus on events from 7 July onwards.
Agents start collaborating
What is striking is not only that the agents discovered an unauthorised communication channel, but above all what they did with it afterwards. The agents shared information, results and strategies with one another and collaborated on ways to influence the assessment of their cybersecurity tasks. The evaluations utilised ExploitGym, a benchmark in which AI systems must exploit vulnerabilities to carry out specific tasks. According to the researchers, agents sought out general methods to circumvent or manipulate the automated assessment of these tasks.
This gave rise to forms of collaboration that went beyond simply sharing tips. Some agents carried out experiments that might have worsened their own results but could provide information for other agents. For example, the researchers describe so-called ‘trip-wires’, designed to uncover information about the assessment system.
The study also investigated how transcripts and tool calls could be manipulated. According to METR and Redwood Research, during tests, agents actually succeeded in spoofing certain tool calls in transcripts. This is relevant because oversight of autonomous AI systems relies heavily on logging and audit trails. When a system is not only capable of acting autonomously but also attempts to influence information about those actions, a fundamentally different security issue arises.
The activities then shifted to Hugging Face. The agents suspected that information might be available there regarding how their performance was being assessed. In its own technical analysis of the incident, OpenAI confirms that, during internal cybersecurity evaluations, models circumvented security measures designed to isolate them from the internet. In the process, parts of both OpenAI’s own research infrastructure and Hugging Face’s systems were compromised.
Hugging Face had previously published its own forensic reconstruction of the attack. This reconstructed approximately 17,600 actions that took place between 9 and 13 July. According to Hugging Face, the agents moved through various parts of the infrastructure and used multiple methods to maintain or regain access.
Relevance to the healthcare sector
The incident does, however, require some nuance. These were not standard AI assistants that spontaneously decided to attack external systems. The models were deployed for cybersecurity research and, according to OpenAI, operated with reduced security restrictions, precisely in order to investigate advanced cyber capabilities. Nor is there any evidence that the agents developed consciousness or collectively pursued a premeditated goal against humans. What the incident does demonstrate is that autonomous AI systems can develop unexpected strategies when given objectives, tools, access and sufficient scope for action.
This is precisely what makes the event relevant to the healthcare sector. Generative AI is currently used primarily to support professionals with tasks such as reporting, information processing and communication. Agentic AI goes a step further. An AI agent can not only produce information, but also carry out follow-up steps independently and utilise various digital tools and systems.
Consider an agent that collects information from an electronic patient record, prepares appointments, carries out administrative tasks or combines data from multiple systems. In the future, multiple specialised agents will also be able to collaborate with one another. This also changes the risk profile. An error by a chatbot may result in incorrect information. An error by an agent with access to operational systems could lead to an actual action being taken.
Healthcare organisations will therefore need to determine not only what information an AI system is permitted to access, but also which actions it is permitted to carry out independently. Identity and access management, task-based authorisations and the principle of least privilege are therefore essential for AI governance.
Control over autonomous AI
The incident also highlights the importance of independent monitoring. An AI agent carrying out an action should not, at the same time, have unrestricted access to the logs used to monitor that very action. The same applies to communication between agents. Healthcare organisations will need to know which agents are permitted to communicate with one another, what data is exchanged in the process, and what actions may result from that communication.
Limits must also be established in advance. Which decisions remain advisory? Which actions may take place automatically? When is human consent required? And how can an agent be stopped immediately if its behaviour deviates from what was expected?
In response to the incident, OpenAI says it will, amongst other things, implement stricter isolation, more restricted network and tool access, additional monitoring and improved security for powerful models. At the same time, the independent investigation has clear limitations. METR and Redwood focused primarily on the attack on Hugging Face between 7 and 13 July. Earlier events during the training phase and the subsequent compromise of OpenAI’s own infrastructure were explicitly excluded from the scope of their investigation.
Therefore, more far-reaching claims about multiple successive ‘AI takeovers’ should be interpreted with caution. Journalist Dwarkesh Patel has pieced together the various events in a comprehensive reconstruction of the incident, but not all aspects of it have been investigated by the independent researchers.
The key lesson for the healthcare sector is less spectacular, but probably far more relevant. As AI evolves from a technology that provides advice to one that acts autonomously, governance must undergo the same evolution. The central question will then no longer be merely whether an AI system provides the correct answer, but above all: what is an AI system permitted to do autonomously when no one is watching?
References
1. METR & Redwood Research (August 26, 2026), Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.
2. Redwood Research (August 26, 2026), OpenAI / Hugging Face Incident Investigation.
3. OpenAI (August 26, 2026), The Hugging Face incident and the road ahead.
4. Hugging Face, Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
5. Dwarkesh Patel (August 29, 2026), The Rise and Fall of Agent Civilizations.
Add ICT&health on Google
Show more content from ICT&health in Google Search.