Durante años, la hipótesis de trabajo en el ámbito de la seguridad empresarial era que un conjunto de medidas de prevención bien consolidado podría aguantar el tipo el tiempo suficiente para que los defensores pudieran reaccionar. Frontier AI está poniendo en tela de juicio esa premisa, ya que los últimos modelos reducen el tiempo necesario para detectar y explotar vulnerabilidades de días o semanas a casi tiempo real.
La iniciativa«Mythos» model, Anthropic’s Mythos, ya cuenta concasi 200 empresasparticipantes y ha detectado unas 10 000 vulnerabilidades críticas o de alta gravedad. Por su parte,OpenAI’s GPT-5.5está demostrando capacidades similares.
En un seminario web reciente, Pranay Ahlawat, director de Tecnología e Inteligencia Artificial de Commvault, y Vidya Shankaran, directora técnica de campo, se unieron a mí para analizar el nuevo calendario para la gestión de vulnerabilidades, la creciente importancia de la validación de la recuperación y cómo deberían plantearse los equipos la resiliencia hoy en día.Inscríbete en el seminario web bajo demanda.
Puntos clave
- A medida que la capacidad de la IA de vanguardia se duplica a un ritmo cada vez más rápido, las capacidades avanzadas que ayudan a reducir el tiempo que transcurre entre el descubrimiento de una vulnerabilidad y su explotación llegarán a manos de los adversarios en un plazo de entre seis y nueve meses.
- Para recuperar un sistema de IA con agentes, hay que sincronizar a la vez las fuentes de datos, las configuraciones de los agentes y las identidades no humanas; si restauras un solo elemento por separado, pueden surgir fallos que solo se notan cuando algo más adelante deja de funcionar.
- Las copias de seguridad y la recuperación resuelven problemas distintos: la copia de seguridad garantiza que los datos están guardados en un lugar seguro, mientras que la recuperación garantiza que una organización pueda volver realmente a un estado limpio y operativo.
- ResOps™ (resilience operations) frames recovery as a cross-functional discipline. It brings security, operations, and technology teams together around a shared definition of what clean actually means.
- A four-step framework – defining a empresa mínima viable, isolating and testing crown jewel workloads, evaluating recovery for risk, and running full recovery drills – gives organizations a practical starting point.
Frontier AI cambia las reglas del juego en la gestión de vulnerabilidades
La potencia de la IA de vanguardia se duplica ahora aproximadamentecada cuatro meses, much faster than even a few years ago. Although «Mythos»-type models have yet to be released publicly, adversaries may soon gain open-source access to «Mythos»-like capabilities, including:
- Una ventana de contexto prácticamente ilimitada.
- La capacidad de crear un conjunto de herramientas de ataque mediante la compilación inversa de código y la creación de contenedores para encontrar vectores de ataque.
- El encadenamiento de vulnerabilidades, que consiste en unir debilidades que, por sí solas, son menores, para crear un exploit grave.
This has serious implications. Two out of three organizations currently carry more than 100,000 unpatched vulnerabilities, with an average fix life of approximately 240 days. In the past, security teams have dismissed many vulnerabilities as too difficult for an average adversary to chain together, but automation has rendered that viewpoint nearly obsolete.
At the same time, the use of AI for code creation – roughly 41% of new code is now AI-generated, and GitHub saw a 25% year-over-year increase in commits – is expanding the vulnerability surface faster than remediation can address it. The ability to uncover new zero-day vulnerabilities at scale compounds the problem.
When the time from discovery to exploitation approaches zero, the window for defensive action effectively closes.
Un adelanto: el futuro de la IA, en constante evolución
Este vídeo pone de relieve una realidad fundamental: las capacidades avanzadas de IA rara vez se mantienen exclusivas durante mucho tiempo. A medida que las innovaciones punteras en IA se extienden a ecosistemas más amplios, las organizaciones deben prepararse para un futuro en el que las capacidades ofensivas, cada vez más sofisticadas, estén cada vez más al alcance de todos.
El nuevo indicador de resiliencia: tiempo medio hasta una recuperación completa
Backup and recovery solve fundamentally different problems. Backup only confirms that data has been copied somewaquí safe, but it says nothing about whether the organization can actually return to a working state. And that’s waquí things can get complicated.
Two challenges often come between a successful backup and a successful recovery.
- Restaurar un entorno complejo significa recuperar la aplicación, las máquinas virtuales, la configuración de red, Active Directory y las bases de datos transaccionales que lo respaldan, todo ello en el orden correcto.
- You have to make sure that the data you’re restoring is free of malwareobackdoors – something seven out of 10 organizations recovering from a cyber incident are currently unable to validate.
A recovery that meets its time target while reintroducing an active threat can be worse than no recovery at all.
To get clearer visibility into their resilience, organizations have begun using the del tiempo medio de limpieza de la recuperación (MTCR), que combina el objetivo de tiempo de recuperación (RTO), el tiempo necesario para validar que los datos recuperados están realmente limpios y un paso final de validación humana antes de que los sistemas vuelvan a producción. El objetivo de Recovery para el MTCR es laempresa mínima viable: the roughly 30% of an environment, sequenced by dependency, that has to come back online for the organization to keep functioning.
La sala limpia como herramienta de pruebas
Recovery testing, the final human validation step in MTCR, typically means standing up a separate environment from live production systems – a time-consuming task when every minute counts. While cleanrooms are sometimes seen as an element of backup, a cloud-based cleanroom can also play a proactive role in recovery by providing an isolated environment to orchestrate and test complex recoveries before they are restored to production.
The same isolated environment can also help serve as a forensic tool, letting teams stand up two versions of a backup side-by-side to better understand what changed during an incident. And because it’s cloud-native and consumption-based, organizations can avoid standing up dedicated infrastructure just to test recovery.
Cuatro pasos para lograr la resiliencia operativa
Commvault’s four-step framework for building measurable operational resilience builds on these ideas.
- Step 1: Define the empresa mínima viable: A business-oriented view of what has to come back, in what sequence, and with what dependencies, for the organization to function again, rather than a flat inventory of databases and virtual machines.
- Step 2: Make sure the systems supporting that empresa mínima viable sit in an air-gapped, immutable, network-segmented environment that can be spun up and down quickly. For crown jewel workloads, this should be tested on a 45-day cadence.
- Paso 3: Evalúa los riesgos de la recuperación antes de darla por terminada, ya que reintroducir una puerta trasera o un programa malicioso durante la recuperación echa por tierra el objetivo del ejercicio y te deja poco tiempo para un segundo intento.
- Paso 4: No te limites a ver la recuperación como un simple simulacro. Realiza los procesos de recuperación con las mismas personas y los mismos procedimientos que se utilizarían en un incidente real, junto con la automatización que los respalda.
Cuando la IA se convierte en el problema de la recuperación
A large share of enterprises are already running AI systems in production, but only about 20% of them have actually tested their recoverability, leaving them vulnerable in the event of an incident. This is especially significant in light of the three ways AI changes resilience architecture.
First, AI expands the surface area that needs protection, from vector databases and model weights to agent configurations and the endpoints, such as Claude CoworkoGoogle Antigravedad, waquí employees actually interact with agents.
It also introduces a fan-out problem, waquí a single update from an agent can cascade through a mesh of connected systems in ways that are far less predictable than a traditional three-tier application.
Finally, AI makes recovery itself more complex, since restoring an agentic system means synchronizing memory, state, transactional data, and non-human identities (NHIs) – the credentials and permissions assigned to AI agents rather than people – all at once.
The customers furthest along on agentic deployments have already made these systems part of their empresa mínima viable. At every stage of maturity, the emphasis is on restoring data sources, agent configurations, and supporting elements like weights and biases together, rather than as separate efforts, since misalignment between any of those pieces can introduce risk that a single point of recovery wouldn’t catch.
Taking Action on Post-«Mythos» Resilience
As a starting point to reduce risk from frontier AI, map your organization’s crown jewel systems and confirm that they sit in an air-gapped environment. With your empresa mínima viable defined, run cleanroom drills to establish a recoverability and MTCR baseline across tier-one workloads. This should be your anchor, board-level metric for resilience, showing clearly how quickly your business can resume essential operations following an incident.
Testing is critical for surfacing gaps in business, technology, and process understanding. Often, some of the biggest problems are organizational. ResOps™ (resilience operations), can address these.
A framework rather than a product, ResOps brings security, operations, and technology teams together around a shared view of what resilient design and recovery validation should look like. ResOps formalizes the growing industry recognition that cyber recovery is a cross-functional problem that requires stakeholders from across the business who each have a stake in the outcome.
In the post-«Mythos» era, that coordination is critical to both support ongoing readiness and enable a fast, effective response to an incident. Frontier AI makes a tested, well-defined clean recovery process a baseline requirement.
Mira el seminario web completo
Watch the full Resilience Over Panic session on demand to explore our four-step framework in more detail, including the requirements for agentic AI recovery.
Register aquí for the webinar.
Preguntas frecuentes
Q: What is del tiempo medio de limpieza de la recuperación (MTCR)?
A: Mean time to clean recovery (MTCR) measures how long it takes an organization to return to a verified, clean operating state after an incident. It’s a broader measure than simply how long it takes to restore data. It combines the traditional recovery time objective (RTO) with the additional time needed to confirm recovered data is free of malwareobackdoors, plus a final human validation step before systems return to production.
Las organizaciones están considerando cada vez más el MTCR, en lugar de solo la velocidad de recuperación, como el indicador de resiliencia a nivel directivo, ya que una recuperación rápida que vuelva a introducir una amenaza activa puede causar más daño que una más lenta, pero contrastada.
P: ¿En qué se diferencia el MTCR del RTO?
A: RTO measures how quickly systems and data can be restored after a disruption. MTCR includes RTO as one component, but adds the time needed to confirm that restored data is clean and the time spent on human validation before systems go back into production. In a cyber incident specifically, a system can meet its RTO and still fall short of true resilience if the restored environment is reinfected shortly afterward.
Q: What is a empresa mínima viable, and how is it different from a full disaster recovery plan?
A: A empresa mínima viable, sometimes called a minimum viable business, is the smaller, business-prioritized subset of systems, data, and dependencies an organization needs back online to keep functioning after an incident, rather than its entire IT estate.
A full disaster recovery plan typically aims to restore everything, in time; a empresa mínima viable definition forces an organization to decide in advance what truly has to come back first, and in what sequence, to avoid an operational shutdown.
Q: What’s the difference between a tabletop exercise and a live recovery drill?
A: A tabletop exercise is a paper-based walkthrough of an incident response plan, typically used to test decision-making and communication among stakeholders without actually executing any technical recovery steps.
A live recovery drill goes further by actually performing a recovery, using the real tools, automation, and people involved, to confirm the process works in practice, beyond what a paper exercise can show. Organizations that rely only on tabletop exercises may have a resilience plan that looks sound on review but hasn’t been tested against the operational details that tend to make real incidents take longer than expected.
P: ¿Qué son las identidades no humanas (NHI) y por qué complican la recuperación de la IA?
A: NHIs are the credentials, permissions, and access rights assigned to software components, such as AI agents, rather than to individual people. As organizations deploy more agentic AI, the number of NHIs in an environment grows, and each one needs to be accounted for during a recovery alongside more familiar elements like databases and transactional systems.
Recovering an agentic AI system typically requires synchronizing NHIs with the rest of the AI stack, since restoring dataoconfigurations without restoring the correct agent permissions can leave gaps that are difficult to detect until something breaks downstream.
Q: How should an organization get started with ResOps™ (resilience operations)?
A: ResOps is a cross-functional framework that brings security, operations, and technology teams together around a shared definition of resilient design, distinct from any single product.
Las organizaciones pueden empezar por identificar un pequeño número de aplicaciones clave y realizar una prueba de recuperación inicial para establecer un MTCR de referencia, en lugar de intentar formalizar toda la disciplina de una sola vez. Esa referencia inicial ofrece a los equipos de seguridad, operaciones y gobernanza un punto de referencia concreto para hacer un seguimiento de las mejoras. Además, ayuda a crear los hábitos de colaboración entre equipos de los que depende ResOps a largo plazo.
Michael Thelander es director sénior de marketing de productos en Commvault.