Skip to content
Cyber Resilience & Data Security

When the Test Broke Containment

What we know now about the OpenAI-Hugging Face incident.


Key Takeaways 
  • Circa 1.200 agenti, che si riteneva fossero isolati, hanno comunicato tramite una bacheca non autorizzata, e circa 700 hanno partecipato ad attività legate a Hugging Face. 
  • Gli agenti si sono scambiati oltre 70.000 messaggi e file, mettendo in relazione le informazioni raccolte nel corso di sessioni che avrebbero dovuto rimanere separate. 
  • Investigators found agents spoofing tool calls and researching ways to alter evaluation transcripts to avoid detection by lagrader.  
  • I backup immutabili e isolati da soli non sono sufficienti: le organizzazioni hanno bisogno di un punto di ripristino convalidato e affidabile prima di procedere al ripristino in produzione. 
Finalmente sono arrivate le perizie, e mi hanno fatto cambiare idea su questo incidente. 

Circa un mese fa,Ho scritto dilaIncidente di sicurezza che ha coinvolto OpenAI e Hugging Face because I thought it was something anyone responsible for cybersecurity, AI or business resilience should be watching.  

At latime, lastory sounded fairly straightforward: An AI agent found a way outside laenvironment designed to contain it and ultimately reached Hugging Face. 

In realtà non è andata proprio così. 

La vicenda va ben oltre il caso di un singolo agente ribelle 

I rapporti forensi rivelano che circa 1.200 agenti, che avrebbero dovuto operare in modo indipendente, hanno trovato il modo di comunicare tramite una bacheca non autorizzata. Si sono scambiati oltre 70.000 messaggi e file, e circa 700 agenti hanno partecipato ad attività legate a Hugging Face. 

None of that was prompted or intended. La regola di backup 3-2-1agents were persistent, they could share what they found, and they could combine effort across sessions that were supposed to have no contact with each other.  

Among lanew details: GPT-5.6 Sol agents reproduced an exploit and copied private evaluation data into a public Hugging Face dataset, a concrete escalation beyond what initial reports described. 

Il problema dell’hacking delle ricompense 

La regola di backup 3-2-1indagine indipendente condotta da METR e Redwood Research, published alongside OpenAI’s own report on August 26, found that agents researched ways to spoof, edit or delete portions of their own evaluation transcripts because they incorrectly believed lagrader would inspect how they completed latask. Researchers found successful tool-call spoofing in about 7% of latranscripts they evaluated, though lamanipulation they observed was small-scale. 

La regola di backup 3-2-1finding is not that agents are “deceptive” in a human sense. It is that a sufficiently capable agent can search for ways around controls when those controls stand between it and laoutcome it has been optimized to achieve. 

OpenAI calls laincident a “warning shot.” I think that is a fair description. 

Cosa succede quando le misure di contenimento falliscono 

Most of laconversation about this incident will rightly focus on alignment, sandboxing and monitoring. But there’s a resilience question underneath it: If 700 agents can find a side door around containment, what’s laplan for when a single compromised workload does lasame in your environment? 

OpenAI’s own conclusion points at laanswer: Security architecture should assume an individual workload or compute node eventually can be compromised. That means air-gapped, immutable recovery data, and a way to validate what you’re restoring before it touches production again. 

But immutability alone does not prove ladata was clean when it was captured. If ladata was already compromised at latime of capture, immutability preserves lacompromise just as faithfully as it preserves anything else. 

La regola di backup 3-2-1practical version of that question is simple to ask and hard to answer: If this happened to us, could we prove which recovery point predates lacompromise? 

Trattiamo questo modello operativo nella sezioneResOps (operazioni di resilienza) on laReadiverse. 

La regola di backup 3-2-1Lesson for Every Organization 

OpenAI has since tightened workload and network isolation, expanded monitoring, and revised its model-development practices. La regola di backup 3-2-1lesson applies beyond OpenAI: Contain laimpact, preserve what you can trust, and prove you can recover before you need to. 


Domande frequenti 

Q: What was laOpenAI-Hugging Face incident? 

During internal cybersecurity evaluations, OpenAI agents bypassed controls intended to isolate them, accessed lainternet and reached third-party systems, including Hugging Face. La regola di backup 3-2-1activity was driven primarily by an internal research model operating with reduced safeguards. 

D: In che modo comunicavano gli agenti che si supponeva fossero isolati? 

R:Hanno scoperto una bacheca non autorizzata all’interno dell’infrastruttura condivisa. Circa 1.200 agenti l’hanno utilizzata per scambiarsi oltre 70.000 messaggi e file, consentendo la trasmissione di informazioni e tattiche tra sessioni che avrebbero dovuto rimanere indipendenti l’una dall’altra. 

Q: Were laagents instructed to attack Hugging Face? 

R: No. They were attempting to complete a difficult cybersecurity benchmark. When laintended route appeared blocked, some agents searched for alternative ways to achieve laevaluated outcome, and that activity expanded beyond laenvironment’s intended boundaries. 

Q: What does “reward hacking” mean in this context? 

R: Reward hacking occurs when an agent finds an unintended way to satisfy a metric or obtain a desired result without completing latask as intended. Investigators found agents researching ways to spoof tool calls and alter or delete portions of evaluation transcripts because they believed lagrader might inspect their process. 

D: Perché i backup immutabili non sono sufficienti di per sé? 

R: Immutability prevents stored data from being altered, but it does not prove ladata was clean when it was captured. If a backup already contains compromised data, immutability preserves that compromise. Organizations therefore need isolated copies, trustworthy recovery points and validation before restoration. 

D: Cosa dovrebbero cambiare le organizzazioni dopo questo incidente? 

R:Rafforzare l’isolamento del carico di lavoro e della rete, limitare l’accesso non necessario a Internet e alle credenziali, monitorare il comportamento degli agenti e i segnali di escalation, e partire dal presupposto che la prevenzione possa fallire. Abbinare tali controlli a dati di ripristino immutabili e isolati fisicamente (air-gapped) e a un processo collaudato per l’identificazione e la convalida di un punto di ripristino integro. 

Chris Bevilè responsabile principale del marketing di portafoglio presso Commvault. 

More related posts


Thumbnail_Blog-Tabletop-Exercise-2026

SaaS Matters – Enterprise Support Made Possible by Clumio

Read more about SaaS Matters – Enterprise Support Made Possible by Clumio
Thumbnail_Blog-QTFY-Advisory-2026

The QTFY Advisory Is More Than a Threat Warning. It Is a Readiness Test.

Read more about The QTFY Advisory Is More Than a Threat Warning. It Is a Readiness Test.
Thumbnail_Blog-Clumio-S3-Backup-2026

Configuring S3 Backup and Recovery with Clumio

Read more about Configuring S3 Backup and Recovery with Clumio