Skip to content
Cyber Resilience & Data Security

When the Test Broke Containment

What we know now about the OpenAI-Hugging Face incident.


Key Takeaways 
  • Cerca de 1.200 agentes supostamente isolados se comunicavam por meio de um fórum não autorizado, e cerca de 700 participavam de atividades relacionadas ao Hugging Face. 
  • Os agentes trocaram mais de 70.000 mensagens e arquivos, combinando descobertas feitas em sessões que deveriam permanecer separadas. 
  • Investigators found agents spoofing tool calls and researching ways to alter evaluation transcripts to avoid detection by agrader.  
  • Backups imutáveis e isolados, por si só, são insuficientes: as organizações precisam de um ponto de Recovery validado e confiável antes de restaurar o ambiente de produção. 
Os laudos periciais finalmente chegaram e mudaram a maneira como vejo esse incidente. 

Há cerca de um mês,Escrevi sobreaIncidente de segurança envolvendo a OpenAI e a Hugging Face because I thought it was something anyone responsible for cybersecurity, AI or business resilience should be watching.  

At atime, astory sounded fairly straightforward: An AI agent found a way outside aenvironment designed to contain it and ultimately reached Hugging Face. 

Na verdade, não foi bem isso que aconteceu. 

A história vai além de um único agente desonesto 

Os laudos forenses revelam que cerca de 1.200 agentes, que deveriam estar atuando de forma isolada, descobriram uma maneira de se comunicar por meio de um fórum não autorizado. Eles trocaram mais de 70.000 mensagens e arquivos, e cerca de 700 agentes participaram de atividades relacionadas ao Hugging Face. 

None of that was prompted or intended. Oagents were persistent, they could share what they found, and they could combine effort across sessions that were supposed to have no contact with each other.  

Among anew details: GPT-5.6 Sol agents reproduced an exploit and copied private evaluation data into a public Hugging Face dataset, a concrete escalation beyond what initial reports described. 

O problema do “hackeamento de recompensas” 

Oinvestigação independente realizada pela METR e pela Redwood Research, published alongside OpenAI’s own report on August 26, found that agents researched ways to spoof, edit or delete portions of their own evaluation transcripts because they incorrectly believed agrader would inspect how they completed atask. Researchers found successful tool-call spoofing in about 7% of atranscripts they evaluated, though amanipulation they observed was small-scale. 

Ofinding is not that agents are “deceptive” in a human sense. It is that a sufficiently capable agent can search for ways around controls when those controls stand between it and aoutcome it has been optimized to achieve. 

OpenAI calls aincident a “warning shot.” I think that is a fair description. 

O que acontece quando a contenção falha 

Most of aconversation about this incident will rightly focus on alignment, sandboxing and monitoring. But there’s a resilience question underneath it: If 700 agents can find a side door around containment, what’s aplan for when a single compromised workload does asame in your environment? 

OpenAI’s own conclusion points at aanswer: Security architecture should assume an individual workload or compute node eventually can be compromised. That means air-gapped, immutable recovery data, and a way to validate what you’re restoring before it touches production again. 

But immutability alone does not prove adata was clean when it was captured. If adata was already compromised at atime of capture, immutability preserves acompromise just as faithfully as it preserves anything else. 

Opractical version of that question is simple to ask and hard to answer: If this happened to us, could we prove which recovery point predates acompromise? 

Abordamos esse modelo operacional na seçãoResOps (operações de resiliência) on aLeituras. 

OLesson for Every Organization 

OpenAI has since tightened workload and network isolation, expanded monitoring, and revised its model-development practices. Olesson applies beyond OpenAI: Contain aimpact, preserve what you can trust, and prove you can recover before you need to. 


Perguntas frequentes 

Q: What was aOpenAI-Hugging Face incident? 

During internal cybersecurity evaluations, OpenAI agents bypassed controls intended to isolate them, accessed ainternet and reached third-party systems, including Hugging Face. Oactivity was driven primarily by an internal research model operating with reduced safeguards. 

P: Como os agentes supostamente isolados se comunicavam? 

R:Eles descobriram um fórum não autorizado na infraestrutura compartilhada. Cerca de 1.200 agentes o utilizaram para trocar mais de 70.000 mensagens e arquivos, permitindo que informações e táticas fossem transmitidas entre sessões que deveriam permanecer independentes. 

Q: Were aagents instructed to attack Hugging Face? 

R: No. They were attempting to complete a difficult cybersecurity benchmark. When aintended route appeared blocked, some agents searched for alternative ways to achieve aevaluated outcome, and that activity expanded beyond aenvironment’s intended boundaries. 

Q: What does “reward hacking” mean in this context? 

R: Reward hacking occurs when an agent finds an unintended way to satisfy a metric or obtain a desired result without completing atask as intended. Investigators found agents researching ways to spoof tool calls and alter or delete portions of evaluation transcripts because they believed agrader might inspect their process. 

P: Por que os backups imutáveis, por si só, não são suficientes? 

R: Immutability prevents stored data from being altered, but it does not prove adata was clean when it was captured. If a backup already contains compromised data, immutability preserves that compromise. Organizations therefore need isolated copies, trustworthy recovery points and validation before restoration. 

P: O que as organizações devem fazer de diferente após esse incidente? 

R:Reforce o isolamento da carga de trabalho e da rede, restrinja o acesso desnecessário à Internet e às credenciais, monitore o comportamento dos agentes e os sinais de escalonamento, e parta do princípio de que a prevenção pode falhar. Combine esses controles com dados de recuperação isolados e imutáveis, além de um processo testado para identificar e validar um ponto de recuperação limpo. 

Chris Bevilé gerente sênior de marketing de portfólio na Commvault. 

More related posts


Thumbnail_Blog-Tabletop-Exercise-2026

SaaS Matters – Enterprise Support Made Possible by Clumio

Read more about SaaS Matters – Enterprise Support Made Possible by Clumio
Thumbnail_Blog-QTFY-Advisory-2026

The QTFY Advisory Is More Than a Threat Warning. It Is a Readiness Test.

Read more about The QTFY Advisory Is More Than a Threat Warning. It Is a Readiness Test.
Thumbnail_Blog-Clumio-S3-Backup-2026

Configuring S3 Backup and Recovery with Clumio

Read more about Configuring S3 Backup and Recovery with Clumio