Per anni, nell’ambito della sicurezza aziendale si è partendo dal presupposto che un sistema di prevenzione ben consolidato potesse tenere testa alle minacce abbastanza a lungo da consentire ai difensori di reagire. Frontier AI sta mettendo in discussione questa premessa, poiché i modelli più recenti riducono i tempi necessari per individuare e sfruttare le vulnerabilità da giorni o settimane a quasi tempo reale.
Lanciata per valutare i potenziali rischi di sicurezza rappresentati dal proprio modelloMythos model, Anthropic’s Progetto Glasswingdi Anthropic ha già coinvoltoquasi 200 aziendee ha portato alla luce circa 10.000 vulnerabilità critiche o ad alta gravità. Nel frattempo,OpenAI’s GPT-5.5sta dimostrando capacità simili.
In un recente webinar, Pranay Ahlawat, Chief Technology and AI Officer di Commvault, e Vidya Shankaran, Field CTO, si sono uniti a me per esplorare le nuove tempistiche nella gestione delle vulnerabilità, la crescente importanza della convalida di Recovery e il modo in cui i team dovrebbero concepire la resilienza oggi.Registrati al webinar on-demand.
Punti di forza
- Man mano che le capacità all’avanguardia dell’IA raddoppiano a un ritmo sempre più accelerato, le funzionalità avanzate che contribuiscono a ridurre il tempo che intercorre tra l’individuazione di una vulnerabilità e il suo sfruttamento saranno a disposizione degli avversari entro sei-nove mesi.
- Il ripristino di un sistema di IA agentica richiede la sincronizzazione simultanea di fonti di dati, configurazioni degli agenti e identità non umane; il ripristino di un singolo elemento in modo isolato può creare lacune che emergono solo quando si verifica un guasto a valle.
- Backup and Recovery risolvono problemi diversi: il backup conferma che i dati esistono in un luogo sicuro, mentre Recovery conferma che un’organizzazione possa effettivamente tornare a uno stato pulito e funzionante.
- ResOps™ (resilience operations) frames recovery as a cross-functional discipline. It brings security, operations, and technology teams together around a shared definition of what clean actually means.
- A four-step framework – defining a “minimum viable company, isolating and testing crown jewel workloads, evaluating recovery for risk, and running full recovery drills – gives organizations a practical starting point.
L’IA di frontiera rivoluziona la gestione delle vulnerabilità
La potenza dell’IA all’avanguardia raddoppia ora all’incircaogni quattro mesi, molto più rapidamente rispetto a pochi anni fa. Sebbene i modelli di tipo Mythos non siano ancora stati resi pubblici, gli avversari potrebbero presto ottenere accesso open-source a funzionalità simili a quelle di Mythos, tra cui:
- Una finestra di contesto praticamente illimitata.
- La capacità di costruire un framework di attacco tramite la decompilazione del codice e la creazione di container per individuare vettori di attacco.
- Il concatenamento delle vulnerabilità, ovvero il collegamento di debolezze singolarmente minori in un exploit grave.
This has serious implications. Two out of three organizations currently carry more than 100,000 unpatched vulnerabilities, with an average fix life of approximately 240 days. In the past, security teams have dismissed many vulnerabilities as too difficult for an average adversary to chain together, but automation has rendered that viewpoint nearly obsolete.
At the same time, the use of AI for code creation – roughly 41% of new code is now AI-generated, and GitHub saw a 25% year-over-year increase in commits – is expanding the vulnerability surface faster than remediation can address it. The ability to uncover new zero-day vulnerabilities at scale compounds the problem.
When the time from discovery to exploitation approaches zero, the window for defensive action effectively closes.
Anteprima: il futuro in rapida evoluzione dell’IA
Questo video mette in luce una realtà fondamentale: le funzionalità avanzate dell’IA raramente rimangono esclusive a lungo. Man mano che le innovazioni all’avanguardia nel campo dell’IA si diffondono in ecosistemi più ampi, le organizzazioni devono prepararsi a un futuro in cui capacità offensive sempre più sofisticate diventeranno più ampiamente disponibili.
Il nuovo indicatore di resilienza: il tempo medio per la Recovery completa
Backup and recovery solve fundamentally different problems. Backup only confirms that data has been copied somewqui safe, but it says nothing about whether the organization can actually return to a working state. And that’s wqui things can get complicated.
Two challenges often come between a successful backup and a successful recovery.
- Il ripristino di un ambiente complesso implica il recupero dell’applicazione, delle macchine virtuali, della configurazione di rete, di Active Directory e dei database transazionali che lo supportano, il tutto nella sequenza corretta.
- You have to make sure that the data you’re restoring is free of malwareobackdoors – something seven out of 10 organizations recovering from a cyber incident are currently unable to validate.
A recovery that meets its time target while reintroducing an active threat can be worse than no recovery at all.
To get clearer visibility into their resilience, organizations have begun using the MTCR (Mean Time to Clean Recovery), che combina l’obiettivo di tempo di ripristino (RTO), il tempo necessario per verificare che i dati ripristinati siano effettivamente puliti e una fase finale di verifica manuale prima che i sistemi tornino in produzione. L’obiettivo di Recovery per l’MTCR è la“minimum viable company: the roughly 30% of an environment, sequenced by dependency, that has to come back online for the organization to keep functioning.
La “cleanroom” come strumento di test
Recovery testing, the final human validation step in MTCR, typically means standing up a separate environment from live production systems – a time-consuming task when every minute counts. While cleanrooms are sometimes seen as an element of backup, a cloud-based cleanroom can also play a proactive role in recovery by providing an isolated environment to orchestrate and test complex recoveries before they are restored to production.
The same isolated environment can also help serve as a forensic tool, letting teams stand up two versions of a backup side-by-side to better understand what changed during an incident. And because it’s cloud-native and consumption-based, organizations can avoid standing up dedicated infrastructure just to test recovery.
Quattro passaggi verso la resilienza operativa
Commvault’s four-step framework for building measurable operational resilience builds on these ideas.
- Step 1: Define the “minimum viable company: A business-oriented view of what has to come back, in what sequence, and with what dependencies, for the organization to function again, rather than a flat inventory of databases and virtual machines.
- Step 2: Make sure the systems supporting that “minimum viable company sit in an air-gapped, immutable, network-segmented environment that can be spun up and down quickly. For crown jewel workloads, this should be tested on a 45-day cadence.
- Fase 3: Valutare i rischi legati alla Recovery prima di dichiararla completata, poiché reintrodurre una backdoor o un malware durante la Recovery vanifica lo scopo dell’operazione e lascia poco tempo per un secondo tentativo.
- Fase 4: Considerare la Recovery come qualcosa di più di una semplice simulazione teorica. Eseguire le operazioni di Recovery avvalendosi delle stesse persone e delle stesse procedure che verrebbero coinvolte in un incidente reale, insieme ai processi automatizzati che ne stanno alla base.
Quando l’IA diventa il problema della Recovery
A large share of enterprises are already running AI systems in production, but only about 20% of them have actually tested their recoverability, leaving them vulnerable in the event of an incident. This is especially significant in light of the three ways AI changes resilience architecture.
First, AI expands the surface area that needs protection, from vector databases and model weights to agent configurations and the endpoints, such as Claude CoworkoGoogle Antigravità, wqui employees actually interact with agents.
It also introduces a fan-out problem, wqui a single update from an agent can cascade through a mesh of connected systems in ways that are far less predictable than a traditional three-tier application.
Finally, AI makes recovery itself more complex, since restoring an agentic system means synchronizing memory, state, transactional data, and non-human identities (NHIs) – the credentials and permissions assigned to AI agents rather than people – all at once.
The customers furthest along on agentic deployments have already made these systems part of their “minimum viable company. At every stage of maturity, the emphasis is on restoring data sources, agent configurations, and supporting elements like weights and biases together, rather than as separate efforts, since misalignment between any of those pieces can introduce risk that a single point of recovery wouldn’t catch.
Agire sulla resilienza post-Mythos
As a starting point to reduce risk from frontier AI, map your organization’s crown jewel systems and confirm that they sit in an air-gapped environment. With your “minimum viable company defined, run cleanroom drills to establish a recoverability and MTCR baseline across tier-one workloads. This should be your anchor, board-level metric for resilience, showing clearly how quickly your business can resume essential operations following an incident.
Testing is critical for surfacing gaps in business, technology, and process understanding. Often, some of the biggest problems are organizational. ResOps™ (resilience operations), can address these.
A framework rather than a product, ResOps brings security, operations, and technology teams together around a shared view of what resilient design and recovery validation should look like. ResOps formalizes the growing industry recognition that cyber recovery is a cross-functional problem that requires stakeholders from across the business who each have a stake in the outcome.
In the post-Mythos era, that coordination is critical to both support ongoing readiness and enable a fast, effective response to an incident. Frontier AI makes a tested, well-defined clean recovery process a baseline requirement.
Guarda il webinar completo
Watch the full Resilience Over Panic session on demand to explore our four-step framework in more detail, including the requirements for agentic AI recovery.
Register qui for the webinar.
Domande frequenti
Q: What is MTCR (Mean Time to Clean Recovery)?
A: Mean time to clean recovery (MTCR) measures how long it takes an organization to return to a verified, clean operating state after an incident. It’s a broader measure than simply how long it takes to restore data. It combines the traditional recovery time objective (RTO) with the additional time needed to confirm recovered data is free of malwareobackdoors, plus a final human validation step before systems return to production.
Le organizzazioni considerano sempre più spesso l’MTCR, piuttosto che la sola velocità di Recovery, come la metrica di resilienza a livello dirigenziale, poiché un Recovery rapido che reintroduca una minaccia attiva può causare danni maggiori rispetto a uno più lento, ma verificato.
D: In che modo l’MTCR si differenzia dall’RTO?
A: RTO measures how quickly systems and data can be restored after a disruption. MTCR includes RTO as one component, but adds the time needed to confirm that restored data is clean and the time spent on human validation before systems go back into production. In a cyber incident specifically, a system can meet its RTO and still fall short of true resilience if the restored environment is reinfected shortly afterward.
Q: What is a “minimum viable company, and how is it different from a full disaster recovery plan?
A: A “minimum viable company, sometimes called a minimum viable business, is the smaller, business-prioritized subset of systems, data, and dependencies an organization needs back online to keep functioning after an incident, rather than its entire IT estate.
A full disaster recovery plan typically aims to restore everything, in time; a “minimum viable company definition forces an organization to decide in advance what truly has to come back first, and in what sequence, to avoid an operational shutdown.
Q: What’s the difference between a tabletop exercise and a live recovery drill?
A: A tabletop exercise is a paper-based walkthrough of an incident response plan, typically used to test decision-making and communication among stakeholders without actually executing any technical recovery steps.
A live recovery drill goes further by actually performing a recovery, using the real tools, automation, and people involved, to confirm the process works in practice, beyond what a paper exercise can show. Organizations that rely only on tabletop exercises may have a resilience plan that looks sound on review but hasn’t been tested against the operational details that tend to make real incidents take longer than expected.
D: Cosa sono le identità non umane (NHI) e perché complicano la Recovery dell’IA?
A: NHIs are the credentials, permissions, and access rights assigned to software components, such as AI agents, rather than to individual people. As organizations deploy more agentic AI, the number of NHIs in an environment grows, and each one needs to be accounted for during a recovery alongside more familiar elements like databases and transactional systems.
Recovering an agentic AI system typically requires synchronizing NHIs with the rest of the AI stack, since restoring dataoconfigurations without restoring the correct agent permissions can leave gaps that are difficult to detect until something breaks downstream.
Q: How should an organization get started with ResOps™ (resilience operations)?
A: ResOps is a cross-functional framework that brings security, operations, and technology teams together around a shared definition of resilient design, distinct from any single product.
Le organizzazioni possono iniziare identificando un numero limitato di applicazioni fondamentali ed eseguendo un test di Recovery iniziale per stabilire un MTCR di riferimento, anziché cercare di formalizzare l’intera disciplina in una sola volta. Tale punto di riferimento iniziale fornisce ai team di sicurezza, operazioni e governance un punto di riferimento concreto per monitorare i miglioramenti. Contribuisce inoltre a sviluppare nel tempo le abitudini trasversali ai team su cui si basa ResOps.
Michael Thelander è direttore senior del marketing di prodotto presso Commvault.