Skip to content

Cloud-Native Backup and Recovery for Google Cloud Storage Now Available.

Clumio for Google Cloud Storage is now generally available – helping to extend immutable backup and rapid recovery to petabyte-scale object storage in Google Cloud.

Key Takeaways

Clumio for Google Cloud Storage is now generally available, helping organizations protect cloud object storage with immutable backups, rapid recovery, and SaaS simplicity. Clumio for Google Cloud Storage helps:

  • Protect Google Cloud Storage data with immutable, air-gapped backups designed to support recovery following ransomware and destructive deletion events.
  • Recover individual objects, prefixes, or entire buckets from a selected point in time.
  • Restore cloud-scale datasets that enable AI, analytics, and business-critical applications.
  • Reduce operational complexity with a fully managed SaaS-based backup and recovery platform.
  • Reduce business disruption with faster recovery workflows.
  • Support compliance and governance initiatives with isolated backup copies and centralized protection policies.

Clumio for Google Cloud Storage is a cloud-native backup and recovery solution that helps deliver immutable, air-gapped protection for Google Cloud Storage objects, prefixes, and buckets at petabyte scale. It helps enable rapid, granular recovery following ransomware, accidental deletion, lifecycle policy errors, or data corruption – without requiring organizations to manage backup infrastructure.

Organizations increasingly rely on Google Cloud Storage as the foundation for analytics platforms, AI initiatives, application data, archives, and cloud-native services.

While Google Cloud provides highly durable storage, durability alone does not address ransomware, accidental deletion, lifecycle policy mistakes, malicious activity, or logical corruption. When cloud object storage becomes a system of record, recovery capabilities become just as important as storage availability.

Clumio for Google Cloud Storage is now available, helping extend cloud-native cyber resilience and recovery capabilities to one of the industry’s leading cloud object storage platforms. The solution delivers immutable, air-gapped backups and rapid recovery workflows that help organizations recover data following cyberattacks, operational mistakes, outages, or corruption events.

As organizations continue investing in AI, analytics, and multi-cloud strategies, the ability to recover large-scale datasets becomes a critical business requirement. Clumio for Google Cloud Storage is designed to help organizations support business continuity, reduce operational overhead, and recover with greater confidence at cloud scale.

Why Google Cloud Storage Requires Dedicated Backup and Recovery

Google Cloud Storage has become the data foundation for modern enterprises. Organizations use it to store training data for AI models, support analytics pipelines, archive business records, and power cloud-native applications.

As data volumes grow, the potential impact of a disruption grows as well. A single lifecycle policy error, accidental deletion, ransomware attack, or application bug can affect millions of objects simultaneously. Native storage durability helps safeguard against infrastructure failures, but it does not address logical corruption, malicious activity, or human error.

The challenge becomes even greater for organizations managing petabytes of data. Recovery efforts often require manual coordination, custom scripts, and time-consuming processes that delay restoration of business-critical services.

Recent research shows that 84% of cloud leaders intentionally use multiple cloud environments. This helps support AI initiatives, balance risk, and manage large datasets. As multi-cloud adoption expands, organizations need consistent resilience and recovery capabilities across environments.

How Clumio Delivers Cloud-Native Resilience

Clumio for Google Cloud Storage was designed to address these challenges through a cloud-native, SaaS-based approach to protection and recovery. The platform stores backup copies in an immutable, air-gapped environment separate from production data.

This isolation helps reduce the risk that backup data is affected if primary storage is compromised by ransomware or destructive deletion events. It helps organizations recover at multiple levels of granularity, including individual objects, prefixes, and entire buckets. This flexibility enables teams to recover only the data they need, helping reduce recovery times and operational disruption.

The solution also helps eliminate the need to deploy, patch, scale, or maintain backup infrastructure. Centralized policy management and automated protection workflows help simplify operations while allowing protection strategies to scale alongside growing cloud data environments.

Business Outcomes for AI, Analytics, and Cloud Operations

For many organizations, downtime is no longer limited to application outages. Disrupted data can halt analytics projects, interrupt AI pipelines, delay customer-facing services, and impact business decision-making. Clumio helps organizations reduce these risks by supporting faster recovery workflows and strengthening cyber resilience. Key business outcomes include:

  • Faster recovery – Rapid point-in-time recovery helps organizations recover data from selected recovery points following ransomware attacks, accidental deletion, corruption events, or lifecycle policy mistakes.
  • Reduced operational overhead – A fully managed SaaS platform that helps remove infrastructure management requirements and reduce reliance on manual recovery processes.
  • Improved cyber resilience – Immutable, air-gapped backups are designed to add an additional layer of resilience against modern cyber threats.
  • Greater confidence in AI and analytics – Organizations can help protect the datasets that enable AI models, business intelligence platforms, and analytics environments.

“With Clumio for Google Cloud, we will be able to restore massive volumes of cloud data with a cloud-native SaaS solution that is easy to use and highly scalable.” – Alex Grach, Head of Engineering, Trusted Data Platform at Atlassian

Bringing Consistent Recovery Across Multi-Cloud Environments

Many enterprises now operate across multiple cloud environments, including AWS and Google Cloud. While cloud adoption creates flexibility, it also can introduce complexity when backup and recovery processes differ between environments.

Clumio helps address this challenge by providing a consistent cloud-native protection experience across cloud platforms. Organizations can apply unified policies, help streamline recovery workflows, and help reduce operational friction as cloud environments grow.

As businesses continue investing in AI-enabled innovation, cloud resilience must evolve alongside the workloads it protects. With the general availability of Clumio for Google Cloud Storage, organizations gain a purpose-built solution for helping protect and recover cloud object storage data at scale while maintaining the operational simplicity that modern enterprises require.


FAQs

Q: What is Clumio for Google Cloud Storage?

A: Clumio for Google Cloud Storage is a cloud-native backup and recovery solution that helps protect Google Cloud Storage data with immutable, air-gapped backups. It helps enable organizations to recover objects, prefixes, and buckets following ransomware, accidental deletion, corruption, or operational mistakes.

Q: Why isn’t Google Cloud Storage durability enough?

A: Google Cloud Storage durability is designed to help safeguard data against infrastructure failures, but it does not address logical corruption, ransomware, malicious deletion, lifecycle policy mistakes, or human error. Independent backup copies help provide an additional recovery layer.

Q: Can Clumio recover individual objects?

A: Yes. Clumio supports granular recovery workflows that allow organizations to recover individual objects, prefixes, or entire buckets depending on the scope of the incident.

Q: How does Clumio help with ransomware recovery?

A: Clumio stores backup copies in immutable, air-gapped environments separate from production storage. These isolated backup copies help organizations recover data following ransomware attacks or destructive deletion events.

Q: Is Clumio designed for large datasets?

A: Yes. Clumio is designed to support cloud-scale environments and can help organizations protect and recover large object storage datasets that enable analytics, AI, and business-critical applications.

Q: How does Clumio simplify operations?

A: As a fully managed SaaS platform, Clumio helps eliminate the need to deploy and maintain backup infrastructure. Centralized policy management and automated workflows help reduce administrative overhead and operational complexity.

Related Resources

PRESS RELEASE 

Clumio Extends Recovery to Google Cloud Storage 

Brings immutable, SaaS-based protection and resilience to petabyte-scale datasets in Google Cloud Storage that are critical for the agentic AI era

Read the announcement  about Clumio Extends Recovery to Google Cloud Storage 
ANALYST REPORT 

The Total Economic Impact of Clumio 

Explore the business value, efficiency gains, and operational benefits organizations achieve with Clumio.

View the report  about The Total Economic Impact of Clumio 
CUSTOMER STORY 

How LoanBoss Strengthens Cloud Resilience 

Learn how organizations strengthen cloud resilience and simplify data protection with Clumio.

Watch the story  about How LoanBoss Strengthens Cloud Resilience 
ANALYST REPORT 

Building Cyber Resilience in the Cloud 

Discover best practices for helping protect cloud-native workloads against ransomware and operational disruptions.  

Read the Report  about Building Cyber Resilience in the Cloud 

Ready to get started?

Resilience Operations

How Resilience Operations (ResOps) Drives Enterprise Recoverability

Enterprise resilience fails not from lack of tools — but because operations, security, and infrastructure teams lack a shared, measurable framework for proving recovery.

It’s 2:47 a.m. and your incident bridge has 40 people on it. The ransomware hit a tier-one workload six hours ago. Containment is done. The forensics team has cleared two recovery points. And now everyone is waiting on the one question nobody prepared for: which services do we restore first, in what order, and how do we know the data is actually clean? Your backup admin pulls up the restore job. Your security lead pulls up the threat report. Your infrastructure lead pulls up the runbook – the one last updated eighteen months ago. Nobody has a shared answer. Nobody has practiced this together. This is the gap that Resilience Operations — ResOps — is designed to close. Not after the incident. Before it.

Resilience Operations (ResOps) is an operational discipline that aligns security, infrastructure, and operations teams around critical services, defined impact tolerances, and continuous validation — so organizations can withstand disruption and demonstrate recoverability with evidence. Commvault Cloud supports ResOps with recovery intelligence, posture visibility, Cleanroom Recovery for isolated restoration, AI-enabled anomaly detection, and automated recovery testing across hybrid, multi-cloud, SaaS, and AI-enabled environments.

Less than 7%

Fewer than 7% of organizations can recover from a ransomware attack within 24 hours of detection. For enterprises running tightly coupled, automated environments — where one compromised workload can cascade into a full-scale operational shutdown — this statistic defines the gap that ResOps is built to close.


What is ResOps and how does it differ from backup and DR?

Backup and disaster recovery are infrastructure disciplines — they answer whether data exists and whether a datacenter can fail over. ResOps is an enterprise operating discipline: it answers whether critical business services can be recovered end-to-end, under real-world stress, within defined impact tolerances — and produces evidence to prove it.

Where DR treats recovery as an IT-owned procedure, ResOps embeds it into the operating rhythm of the entire enterprise. A ResOps Council gives cross-functional teams — engineering, security, infrastructure, operations, service delivery — shared ownership and decision rights over resilience outcomes. Recovery is then governed by two measurable targets: Service Resilience Indicators (SRIs), which define how well each critical service must perform under disruption, and Mean Time to Clean Recovery (MTCR), which tracks how quickly it actually gets there. The result is a shift from annual DR tests and static runbooks to continuously measured, board-reportable recoverability.

  • Critical Services Mapping: ResOps begins by identifying the Minimum Viable Company (MVC) — the smallest set of critical services required to sustain business operations — and defining acceptable impact tolerances for each. This scope definition informs downstream governance and testing priorities.
  • SRI-Based Performance Targets: Each critical service is assigned a Service Resilience Indicator (SRI) — a specific, testable target for how the service should perform under disruption. SRIs replace vague recovery intentions with more accountable, measurable targets.
  • MTCR Tracking: Mean Time to Clean Recovery (MTCR) measures the elapsed time from incident declaration to verified restoration of a critical service. Unlike recovery time objective (RTO), which measures uptime restoration, MTCR incorporates validation steps to support recovery confidence.
  • Cross-functional Decision Rights: ResOps defines who makes recovery decisions, in what order, and under what conditions — using RACI diagrams, backup authority structures, and preapproved runbooks. This helps reduce coordination gaps that can occur when siloed teams respond during an incident.
  • Continuous Validation Cadence: ResOps replaces annual disaster recovery (DR) tests with an ongoing rhythm of simulations, tabletop exercises, and cleanroom restores — each producing evidence that recovery capabilities remain current and effective.

The ResOps framework: Five integrated domains for enterprise resilience

The ResOps framework is a closed-loop operational model built around five integrated domains — Resilience Governance, Recovery Planning, Recovery Architecture, Resilience Assurance, and Resilience Measurements —that together help organizations maintain critical services within defined impact tolerances during disruption. Once these domains are established, teams can adopt a posture of continuous improvement, transforming resilience from a one-time project into an ongoing, measurable program.

Each domain plays a specific role in the ResOps closed loop. Resilience Governance establishes the charter, defines the Minimum Viable Company (MVC), and creates cross-functional ownership through a ResOps Council. Recovery Planning and Recovery Architecture translate that governance into testable runbooks, recoverability tiers, separation of control and data planes, immutable recovery points, and air-gapped isolation. Resilience Assurance validates these architectures through continuous testing — such as simulations, cleanroom restores, and tabletop exercises — while Resilience Measurements track SRIs, MTCR, and reporting outputs to support visibility and decision-making.

  • Resilience Governance: Establishes a ResOps charter, defines impact tolerances for each critical service, aligns resilience outcomes to organizational funding, and creates a cross-functional ResOps Council for shared accountability and decision-making.
  • Recovery Planning: Defines recoverability tiers by business criticality, including technical runbooks for system restart, RACI diagrams of recovery responsibilities, dependency maps, and testing schedules — so teams have a documented and rehearsed path to recovery.
  • Recovery Architecture: Defines separation between control planes, data planes, and storage tiers; incorporates air-gapping, immutability, and domain isolation; and is designed to reduce blast radius within the recovery environment to support faster, more controlled restoration.
  • Resilience Assurance: Embeds continuous validation into the operating rhythm — including cleanroom restores, simulations, and governed tabletop exercises —designed to validate recovery processes and reduce reinfection risk.
  • Resilience Measurements: Tracks outcome-focused metrics, including impact tolerances, MTCR, SRI attainment, and critical service status, and translates them into quarterly resilience reporting for leadership visibility.

How Commvault Cloud supports ResOps across hybrid enterprise environments

Commvault Cloud supports ResOps as an evidence-driven platform that helps extend cyber resilience beyond protection tooling into a broader operational model. While ResOps is a discipline rather than a product, Commvault Cloud provides capabilities that help organizations operationalize its five domains across hybrid, multi-cloud, SaaS, and AI-enabled environments.

Commvault Cloud helps address the ResOps evidence gap by combining recovery intelligence, posture visibility, and recovery workflows within a unified platform. Rather than stitching together point tools for backup, disaster recovery (DR), and security operations, Commvault Cloud provides a unified data protection console across enterprise workloads — on-premises, cloud, and SaaS —while integrating with SecOps tools so detection and recovery can operate in a coordinated manner. This approach enables cross-functional teams to define SRIs, measure MTCR, and continuously validate recovery processes in isolated environments — supporting the evidence needs of stakeholders, including leadership and regulators.

  • Cleanroom Recovery: Commvault’s Cleanroom Recovery capability provisions an isolated, air-gapped environment on demand — separate from the production network— where critical workloads can be restored, analyzed, and scanned prior to returning services to production. (See FAQ Q3 for step-by-step mechanics.)
  • AI-Enabled Anomaly Detection: Commvault Cloud’s AI-enabled detection helps identify unusual data access patterns and backup anomalies early — helping limit incident impact and reduce the scope of recovery.
  • Automated Recovery Testing: Instead of relying solely on periodic DR exercises, Commvault Cloud supports ongoing validation of recoverability by running non-disruptive recovery tests and comparing results against SRI targets, helping surface gaps before an incident.
  • Unified Data Protection Console: A single control plane spans on-premises, cloud, SaaS, and AI-enabled workloads — helping reduce coverage gaps and tool sprawl that can affect recovery confidence in hybrid environments.
  • Posture Visibility and SRI Reporting: Commvault Cloud provides visibility into resilience posture across protected workloads in a unified view — tracking SRI attainment, MTCR trends, and tolerance gaps—and generating reporting artifacts to support internal and external stakeholders.

Microsoft Sentinel (SIEM)

Bidirectional integration allows Commvault Cloud to pass recovery telemetry into Microsoft Sentinel for correlation with threat detections — helping inform recovery decisions with current security context.

CrowdStrike Falcon (Security Platform)

Integration with CrowdStrike provides threat intelligence that helps inform recovery point selection — so restored environments can be assessed against known indicators of compromise before returning to production.

Splunk (SIEM/SOAR)

Commvault Cloud sends recovery event data to Splunk to provide unified visibility across security operations, helping teams correlate backup anomalies with broader threat activity.

Microsoft Azure / AWS / Google Cloud (Cloud)

Commvault Cloud’s any-to-any workload portability supports recovery across major hyperscalers — helping organizations maintain resilience as dependencies shift across cloud environments.

ServiceNow (ITSM)

Integration with ServiceNow supports automated incident ticket creation and orchestration of recovery workflows — helping connect security detection with IT operations response.

How ResOps works end-to-end: From governance to clean recovery


Discover

The unified data protection console helps classify enterprise data and map service dependencies — creating a centralized inventory of what constitutes the Minimum Viable Company (MVC) and which workloads align to specific recoverability tiers.


Protect

Policy-driven protection is applied across workloads based on recoverability tiers defined in Recovery Planning. This results in recovery points designed with immutability and air-gap principles, reflecting the Recovery Architecture approach — such as separation of control planes, data planes, and storage, and considerations for blast-radius reduction.


Detect

AI-enabled anomaly detection monitors backup telemetry and data access patterns. When irregularities are identified, alerts can be routed to integrated SecOps platforms — helping security and recovery teams operate from a shared signal and reduce delays in response.


Recover

Following incident declaration, orchestrated workflows run against pretested runbooks — reducing the need for ad hoc response. Cleanroom Recovery provisions an isolated environment where recovery points can be analyzed and tested before services are returned to production.


Restore

Services are promoted to production based on SRI priority. MTCR is captured for each service. The recovery sequence — including timestamps, validation steps, and SRI attainment — can be logged and compiled into reporting artifacts to support internal reviews and stakeholder reporting.

ResOps transforms enterprise resilience from documentation-based planning into a continuously validated, evidence-led operating discipline. Organizations that operationalize ResOps gain a cross-functional framework that unites security, IT, and infrastructure around measurable recoverability — tracked through SRIs, MTCR, and reporting for leadership visibility.

Commvault Cloud supports this model through Cleanroom Recovery, AI-enabled anomaly detection, automated recovery testing, and unified data protection across enterprise workloads.

The result: when disruption occurs — from ransomware, AI-enabled failure, or cascading infrastructure outages — teams are better prepared, recovery processes are validated, and organizations can provide supporting evidence to stakeholders, including regulators.

Frequently Asked Questions

What is ResOps and how does it differ from traditional backup and disaster recovery?

ResOps addresses a gap backup and DR don’t fully cover: can critical services be recovered end-to-end under real-world conditions within defined impact tolerances — and can we prove it? Backup confirms data copies exist, and DR validates data center failover, but ResOps extends this by accounting for dependencies, clean-state validation, rebuild pathways, and cross-functional execution, with metrics like SRIs and MTCR supported by Commvault Cloud.

How does operational resilience differ from business continuity planning (BCP)?

Business continuity planning (BCP) defines how an organization intends to respond to disruption, producing documented procedures that are tested periodically. Operational resilience and ResOps focus on continuously testing and improving an organization’s actual ability to withstand disruption within defined tolerances. Commvault Cloud supports this shift by providing the measurement and validation layer — SRI tracking, MTCR reporting, and Cleanroom-based testing — that enables organisations to demonstrate execution rather than simply document intent.

How does Commvault Cloud’s Cleanroom Recovery work for ransomware response?

When an incident occurs, Commvault Cloud provisions a Cleanroom environment — an isolated network segment designed to limit exposure to compromised systems. Recovery points are scanned for potential threats before validation activities — such as application startup and dependency checks— help confirm readiness, with the process generating logs and artifacts that can support internal review and regulatory reporting.

How is ResOps different from what Rubrik, Cohesity, or Veeam offer?

Rubrik, Cohesity, and Veeam provide data protection and recovery capabilities, while ResOps introduces an operating model that aligns security, operations, and infrastructure teams around defined impact tolerances and evidence-based recoverability. Commvault Cloud supports this approach with capabilities such as a unified control plane, Cleanroom Recovery, AI-enabled anomaly detection, and automated measurement of resilience metrics like SRIs and MTCR.

Which compliance frameworks require operational resilience evidence, and how does ResOps address them?

Frameworks such as NIS2, the EU Cyber Resilience Act, DORA, and NIST CSF 2.0 emphasize resilience, testing, and accountability, though specific requirements vary by regulation and jurisdiction. ResOps can help organizations align to these expectations by providing an operating model and measurable outputs — such as SRI metrics and recovery validation records— supported by Commvault Cloud capabilities.

When should an enterprise adopt ResOps vs. simply improving its existing DR program?

Organizations may improve DR when addressing specific gaps like recovery time objectives, recovery point objectives, or workload coverage. ResOps becomes relevant when challenges are broader — siloed teams, unclear dependencies, or limited visibility into recovery readiness. Commvault Cloud supports the transition from DR to ResOps by providing a unified control plane, Cleanroom Recovery for validated testing, and SRI-based measurement that makes recovery readiness visible and reportable to leadership and regulators.

Prove Your Resilience with Commvault Cloud ResOps

Start with critical services, define impact tolerances, and validate clean recovery with evidence — using SRIs, MTCR, and Commvault Cloud.

Related resources

Explore

What Is Resilience Operations (ResOps)?

Understand the ResOps operating model — how it unites data security, identity resilience, and cyber recovery into a continuous discipline for AI-era enterprises.
Read the article about What Is Resilience Operations (ResOps)?
Blog

Re-envisioning Resilience for the Age of AI

Discover how to actively manage resilience across increasingly complex AI environments with a new cross-functional operational approach powered by Commvault Cloud.
Read the blog about Re-envisioning Resilience for the Age of AI

To address mandates governing where their data is stored and used, many organizations think they can buy a sovereign cloud SKU from a hyperscaler and check the box. But this falls far short of what’s actually required – something they might discover only when a regulator asks them to demonstrate that a dataset never left a defined geography, that no foreign-jurisdiction personnel accessed it, and that they can recover it within 24 hours under incident conditions.

In a recent webinar, I joined Commvault GM Alex Zinin, who leads our digital sovereignty task force, along with Jakub Lewandowski, our associate general counsel for EMEA, and Pranay Ahlawat, our chief technology and AI officer, to examine what a complete approach to digital sovereignty needs to include and where most programs fall short.

See the full webinar and get the full Digital Sovereignty Decoded readiness report and implementation framework.

Key Takeaways

  • The EU Cloud Sovereignty Framework defines eight sovereignty objectives, only one of which concerns data location.
  • Selecting a sovereign cloud region addresses where data lives, but not who can access it, under what legal authority, or whether you can recover it under real conditions.
  • A complete sovereignty posture spans four interdependent pillars: data locality, technological sovereignty, operational sovereignty, and jurisdictional sovereignty.
  • Sovereignty programs that treat recovery architecture separately from primary data governance carry an unexamined risk that can surface during incidents.
  • Rather than a policy of maximum sovereignty at any cost, organizations should design their strategy around minimum viable sovereignty: the right controls, consistently enforced, calibrated to actual obligations.

Digital Sovereignty Becomes an Enterprise Requirement

Over the past decade, the toughest digital and data sovereignty rules have mainly applied to government, defense, and other national‑security workloads. Outside of highly regulated sectors, many enterprises treated sovereignty principles as design guidance rather than a hard architectural constraint.

That’s now changing. GDPR enforcement has matured beyond guidance into substantial fines for operational failures, and newer regimes like DORA, NIS2, Germany’s KRITIS rules, and the EU Data Act have tightened expectations around jurisdictional control and operational resilience.

Sovereignty questions are now surfacing in RFPs, M&A due diligence, and board‑level risk reviews as well.

The EU Cloud Sovereignty Framework, published in October 2025, clarifies what this scrutiny actually evaluates. Of its eight sovereignty objectives, only one addresses where data resides. The other seven cover access control, operational dependencies, jurisdictional exposure, and recovery.

For enterprises, this structure is now the lens through which vendor capabilities must be evaluated.

Sneak Peek: Rethinking Sovereignty and Resilience

In this moment from the webinar, Jakub discusses how sovereignty is a risk posture, rather than a single product. Organizations need a holistic strategy that combines architecture, operations, governance, auditability, and recovery planning to address it.

Data Residency Does Not Equal Sovereignty

Data residency only answers questions about where. Sovereignty regulations also require being able to explain who, how, and under what conditions.

In practical terms, a complete sovereignty posture encompasses four interdependent pillars.

1. Data Locality

This pillar covers not just where data is stored, but where it travels. Control-plane artifacts, metadata, and telemetry can cross geographic boundaries even when primary data stays in-region.

2. Technological Sovereignty

This pillar addresses whether the organization controls the mechanisms that protect data:

  • How access is granted.
  • How data is encrypted.
  • Whether encryption key custody is retained under all conditions.

A key concept here (pun intended) is the distinction between Bring Your Own Key (BYOK), where an organization’s own encryption keys are managed within the provider’s platform, and Hold Your Own Key (HYOK), where the organization retains independent custody of keys entirely outside the provider’s environment.

For regulated organizations with strict sovereignty requirements, BYOK may not provide sufficient protection if a foreign legal authority can compel the provider to surrender key access under some circumstances.

3. Operational Sovereignty

This covers who operates the environment and from where, including whether support personnel or third-party vendors are subject to foreign jurisdiction.

4. Jurisdictional Sovereignty

This final pillar establishes the legal framework under which services are delivered and whether there are explicit protections against extraterritorial access, such as the cross-border situations discussed in the U.S. CLOUD Act.

Each of these pillars is essential to maintain compliance. A strong data locality posture with weak operational controls can allow unexamined risk.

Where Digital Sovereignty Programs Break Down

My experiences in the field have revealed a recurring pattern: When sovereignty becomes a technical conversation, it gets too narrow, too quickly. Workshops zoom in on where data lives, teams get to work on that one question, and then they move on, leaving the other three pillars largely unexamined.

Structural problems also come into play. Digital sovereignty needs to be treated as an ongoing program with legal, technical, and operational stakeholders, not as an IT project that gets checked off as complete.

Organizations can also be led astray by a few myths. One, as we’ve discussed, is the impression that sovereignty equals residency.

Then there’s the myth of absolute sovereignty, the idea that you can achieve complete independence from all external jurisdictions and dependencies. In practice, sovereignty always involves tradeoffs between control, cost, technological velocity, and the ability to innovate.

The goal should be to strike the right balance between independence from foreign jurisdictions and the requirements of your business. It’s also important to understand that sovereignty isn’t a product you can buy, but a risk posture built from architecture, operations, contracts, certifications, and continuous auditability.

Resilience Belongs Inside the Sovereignty Boundary

Operational sovereignty is the hardest pillar to audit and the one most commonly underestimated. If your environment needed access for routine maintenance tonight, who would perform it, from which country, and under which legal jurisdiction?

Most organizations, when they work through that question for the first time, find at least one support pathway that crosses a jurisdiction boundary they hadn’t mapped.

This gap becomes most consequential during recovery. Most sovereignty programs govern primary data environments but treat backup infrastructure, restoration sequencing, and recovery point management under a separate – and often weaker – set of controls. When an incident occurs, recovery personnel may not meet jurisdictional requirements, and the sovereign architecture designed to protect data can actively complicate restoration if resilience wasn’t designed in from the start.

Commvault’s resilience operations model, ResOps™, addresses this need directly by framing recovery as an ongoing operational discipline that needs to be designed, tested, and validated inside the same sovereignty boundary as the data it protects.

Minimum Viable Sovereignty: The Right Level of Control, Not the Maximum

An absolutist approach to digital sovereignty can tax resources while unnecessarily restricting a company’s ability to meet its business goals.

A payroll system, a customer transaction database, and an internal HR tool don’t carry the same sovereignty obligations. Taking a binary approach to compliance can lead to either under-investing where it matters or over-investing beyond what’s actually required.

Minimum viable sovereignty sets a more practical target: the right controls, consistently enforced and continuously demonstrated, calibrated to what each workload actually requires across all four pillars.

Organizations have a broad range of options for how sovereign controls are delivered across their environment, each providing different controls.

How Commvault Is Addressing the Digital Sovereignty Issue

Commvault’s Geo Shield framework is designed to help organizations navigate that spectrum. Rather than offering a single sovereign SKU, Geo Shield maps to the full spectrum of deployment models:

  • Regional sovereign cloud services delivered as SaaS
  • Hyperscaler sovereign-launch partnerships
  • Partner-operated national sovereign offerings built with local service providers
  • Fully customer-controlled private sovereign environments qualifying under frameworks like FedRAMP High.

In this way, organizations can achieve a digital sovereignty posture that holds up in real-world conditions – even when an incident occurs.

Watch the Full Webinar and Get the Readiness Report

In the full webinar, available on demand, you’ll discover:

  • Why digital sovereignty is more than a technology solution.
  • The role of architecture and operations in sovereignty strategy.
  • How governance, contracts, and auditability impact resilience.
  • Why sovereignty must hold up during cyber incidents and outages.
  • The importance of a holistic, risk-based approach to sovereignty.

Watch the webinar and get the full Digital Sovereignty Decoded readiness report and companion framework.

FAQs

Q: What is the difference between data residency and digital sovereignty?

A: Data residency addresses where data is physically stored. Digital sovereignty is broader: it addresses who can access data, under what legal authority, through which operational pathways, and whether it can be recovered cleanly under real conditions.

An organization can have data residing in the right country while remaining exposed to foreign jurisdiction through its support personnel, vendor access agreements, or backup infrastructure. Residency is the starting condition; sovereignty is the full posture built on top of it.

Q: What is the EU Cloud Sovereignty Framework, and why does it matter?

A: The EU Cloud Sovereignty Framework is a structured assessment tool developed by the European Commission to evaluate cloud and technology providers against sovereignty criteria during procurement processes.

It defines eight sovereignty objectives, with assurance levels ranging from zero to four for each. Only one of the eight objectives addresses data location; the rest cover operational controls, key custody, jurisdictional exposure, and recovery.

The framework represents the most comprehensive public framework for evaluating sovereignty posture and is increasingly being used as a reference by other regions and procurement bodies beyond the EU.

Q: What is the difference between BYOK and HYOK, and why does it matter for sovereignty?

A: Bring Your Own Key (BYOK) allows an organization to supply its own encryption keys, but those keys are typically managed within the provider’s platform. Hold Your Own Key (HYOK) means the organization retains independent custody of keys entirely outside the provider’s environment, including under crisis conditions or legal compulsion.

For regulated organizations with strict sovereignty requirements, BYOK may not provide sufficient protection if a foreign legal authority can compel the provider to surrender key access. The U.S. CLOUD Act, for example, can reach providers operating under U.S. jurisdiction regardless of where data is physically stored.

HYOK addresses that exposure directly, though it may require a higher platform tier in SaaS deployments.

Q: Why do most sovereignty strategies overlook operational sovereignty?

A: Operational sovereignty, covering who operates the environment and from where, is the hardest pillar to audit because it requires inventorying support contracts, vendor access agreements, and third-party dependencies across the full operational chain.

Most organizations start sovereignty programs focused on data location and encryption, which are more visible. Operational dependencies tend to surface only when explicitly audited or when an incident forces the question.

As a first step to evaluate operational sovereignty, you should identify every access pathway into your sovereign environment and the legal jurisdiction of each party with that access.

Q: How should organizations think about recovery in the context of sovereignty?

A: Recovery architecture needs to meet the same sovereignty requirements as primary data environments, but it often doesn’t. In most organizations, backup infrastructure, restoration sequencing, and recovery point management are frequently governed by a separate set of controls, or none at all.

During an incident, the personnel authorized to execute recovery may not meet jurisdictional requirements, recovery points may not have been validated as clean and uncompromised, and the sovereign architecture designed to protect data can actively complicate recovery if resilience wasn’t built into the original design. A sovereignty review should always include recovery planning.

Q: What does minimum viable sovereignty mean in practice?

A: Minimum viable sovereignty means identifying the right level of control for each workload, calibrated to actual regulatory obligations, risk tolerance, and operational constraints, rather than applying maximum controls uniformly.

Maximum sovereignty comes with real tradeoffs: technological complexity, operational burden, service limitations, and cost. Organizations that define requirements by workload across the four pillars, map those requirements to deployment models, and build evidence of consistent enforcement are in a far stronger position than those pursuing all-or-nothing approaches.

Q: What role do certifications like C5, SecNumCloud, and ISO 27001 play in a sovereignty strategy?

A: Certifications help provide auditable evidence that controls have been independently verified, an important part of any defensible sovereignty posture. C5 in Germany, SecNumCloud in France, and ISO/IEC 27001 each establish baseline requirements that providers must demonstrate through independent audits.

The strongest sovereignty postures treat these certifications as a floor, providing necessary evidence that controls exist, but not a substitute for operational testing under realistic conditions.

Darren Thomson is Vice President and Chief Technology Officer, EMEA, at Commvault. Be sure to catch him in the podcast series, STRIVE.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

Key Takeaways

  • Multiple backup tools often create operational dependencies on a small number of specialists, increasing organizational risk.
  • Managing protection across separate consoles, policies, and reporting systems makes it harder to maintain visibility and respond quickly to issues.
  • Consolidation does not have to mean a disruptive rip-and-replace project; many organizations can modernize gradually while retaining existing infrastructure investments.
  • A unified control plane can help simplify policy management, monitoring, auditing, and recovery operations across hybrid environments.
  • Organizations that simplify backup often realize significant cost savings while helping improve operational efficiency and resilience.

You didn’t build a messy environment. You built a functional one.

Each tool in your backup stack solved a real problem when you added it. One handled virtual machines. Another covered cloud workloads. A third came in when the business moved to SaaS. You made smart calls with the budget and the vendors you had. The environment works.

The stack isn’t the problem. It’s the operational model that comes with it.

Today, it’s not uncommon for a data center team managing legacy backup infrastructure to run seven or more separate systems. Seven sets of policies. Seven consoles. Seven renewal cycles.

And, in the background, they also get seven single points of failure: not in the infrastructure, but in the people. Because somewhere in your organization, there are one or two engineers who know how each of these systems behaves. When something breaks at 2 a.m., you know exactly who’s getting the call.

That’s not resilience. That’s dependency masquerading as expertise.

The Dashboard Wall

Here’s a question worth sitting with: How long does it take your team to answer a simple question like “Did last night’s backup run clean across all workloads?”

If the answer involves opening more than one console, you already know the problem. Each tool has its own view of the world. Each one reports on what it protects, in its own format, on its own schedule.

Stitching that picture together – across on-premises systems, cloud workloads, and remote locations – takes time your team doesn’t have and creates gaps that only show up when something goes wrong.

The scripts help. Your team probably wrote them. But scripts that bridge what tools don’t natively share are technical debt with a support contract. They work until they don’t, and when they don’t, the fix requires the person who wrote them.

And if you want to do this across AI-dependent workloads using generated data … let’s just say you increased the degree-of-difficulty factor by 100% or more.

What Consolidation Really Means for Infrastructure Teams

The instinct when you hear “consolidate your backup environment” is to picture a rip-and-replace project with new hardware, new procurement, and a migration that takes six months while landing at the worst possible time.

But that’s not what consolidation has to look like.

The right platform works with the storage already in your rack. It doesn’t require you to throw out contracts you negotiated or hardware you haven’t depreciated. You can start where it makes sense – remote offices, a specific cloud workload, a dataset that’s been a problem – and expand as old contracts run out and budget frees up.

What you get in return is a single control plane. One place to set policy, monitor protection, and answer the auditor’s question. One operating model that works across on-premises, cloud, and hybrid workloads without a script to bridge the gap.

The engineers who were keeping seven dashboards cobbled together through scripts and custom executables start doing something more useful instead.

The Proof Is in the Number

Fortune Brands consolidated its backup environment with Commvault® Cloud and saved $22.7M – a 73% reduction in total cost. NTT-Netmagic reduced costs by $300K annually and cut storage overhead by 35%.

Those aren’t modernization-project numbers. They’re operational-relief numbers. The kind that come from stopping the compounding cost of complexity – not from buying new things.

Real Resilience Doesn’t Need a War Room

If running a recovery drill requires assembling a team of specialists who each know one piece of the environment, that’s not a drill. That’s a liability.

Real resilience means any qualified engineer on your team can execute recovery. It means one set of policies, one control plane, and a recovery process that doesn’t fall apart when the person who built it is on vacation.

Seven dashboards can protect your data. They can’t protect your team from the operational weight of keeping them running.

That’s the case for consolidation. Not a better product. A better way to run what you’ve built.

Want the full picture? Download The Hidden Cost of Seven Tools – a field guide for data center teams who built something worth protecting.

FAQs

Q: Why is managing multiple backup platforms a problem if they’re all working?

A: The challenge isn’t usually whether the tools function individually – it’s the operational burden of managing them together. Multiple consoles, policies, and reporting systems can make visibility, troubleshooting, and recovery more complex than they need to be.

Q: What is one of the biggest risks created by a fragmented backup environment?

A: In many organizations, critical knowledge becomes concentrated in a few individuals who understand how specific systems interact. If those team members are unavailable during an incident, recovery efforts can become slower and more difficult.

Q: Does consolidation mean replacing all existing infrastructure?

A: Not necessarily. Many consolidation initiatives are phased approaches that work alongside existing storage, hardware, and contracts. Teams can modernize gradually based on business priorities, budget cycles, and contract renewals.

Q: How can consolidation improve resilience?

A: A unified platform can help provide consistent policies, centralized visibility, and streamlined recovery processes. This helps enable more team members to confidently execute recovery procedures without relying on specialized knowledge tied to individual tools.

Q: What about vendor lock-in when consolidating to a single platform?

A: Vendor lock-in is a valid consideration. The goal of consolidation should be to help reduce operational complexity while maintaining flexibility through open architectures, broad workload support, and the ability to leverage existing infrastructure investments where possible.

Q: How do organizations measure the value of consolidation?

A: Beyond software costs, organizations often evaluate factors such as administrative overhead, recovery efficiency, storage utilization, training requirements, audit readiness, and the reduction of operational risk. The greatest value frequently comes from simplifying day-to-day operations and improving recovery confidence.

Michael Thelander is Senior Director, Product Marketing, at Commvault.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

From Detection to Recovery: What Does Modern Cyber Resilience Architecture Require?

Modern cyber resilience helps enable organizations to restore trusted operations after compromise using validated recovery, isolated environments, and coordinated response across hybrid infrastructure.

Key Takeaways

Modern cyber resilience focuses on trusted recovery by validating data, isolating restoration, coordinating response, and enabling flexible recovery across hybrid environments.

  • Modern cyber resilience depends on evidence-driven recovery that validates data integrity before restoration, not just backup availability.
  • Traditional disaster recovery models break under ransomware because attackers target backups, extend dwell time, and compromise restore points.
  • A resilience operations (ResOps) model aligns security, IT, and data teams around continuous validation, clean recovery workflows, and measurable readiness.
  • Cyber recovery must integrate with the broader security ecosystem, connecting detection, response, and recovery systems to enable coordinated action and shared visibility during incidents.
  • Cleanroom Recovery®, Synthetic Recovery™, air-gapped backups, and AI-assisted detection work together to help enable trusted, isolated restoration.
  • Workload portability and minimum viable recovery help organizations restore critical business functions first and recover across hybrid environments without platform constraints.

Most organizations can detect cyberattacks. Far fewer can recover cleanly, confidently, and at scale across their entire data estate. Commvault addresses this gap through evidence-driven recovery — combining anomaly detection, Cleanroom Recovery, Synthetic Recovery, and the ResOps operating model to help organizations validate, isolate, and restore trusted operations even under active adversarial conditions.

Why Do Traditional Recovery Models Fail Against Modern Cyberattacks?

241 days. That’s how long the average breach lifecycle is, according to IBM’s Cost of a Data Breach Report 2025. The report also showed that 76% of organizations still took more than 100 days to fully recover from a breach, giving attackers ample time to compromise backup systems and recovery points.

Legacy disaster recovery strategies were designed for outages and hardware failures, not adversarial attacks. They assumed backups could be trusted by default – an assumption that no longer holds.

Cyberattacks are no longer isolated security events. They are enterprise-wide disruptions that expose how well teams can respond and recover amid fragmented tools, signals, and decision-making.

Attackers move patiently and deliberately. By the time encryption or destruction occurs, multiple restore points may already be unsafe.

Today’s ransomware and cyberattacks follow a playbook that can look like this:

  • Extended dwell time: Adversaries may remain in the environment for weeks or months, during which they modify files, insert dormant malware, steal credentials, and corrupt backup repositories.
  • Backup targeting: Attackers now can actively delete snapshots, disable backup jobs, exfiltrate recovery keys, and alter stored data.

By the time encryption occurs, multiple restore points could be compromised.

The IDC MarketScape: Worldwide Cyber-Recovery 2025 Vendor Assessment highlights that modern recovery must enable both data survival and integrity, especially when attackers target protection layers directly.

These conditions expose systemic gaps. Security and recovery teams often operate independently, creating delays in decision-making. Backup systems lack built-in validation, leaving teams uncertain about what is safe to restore. Recovery environments may not be isolated, increasing the risk of reinfection.

Closing these gaps is fundamental to modern cyber resilience. A unified approach that brings together anomaly and threat detection, data protection, AI-assisted insights, and recovery validation is key.


Why Is Evidence-Driven Cyber Recovery the Way Forward?

Evidence-driven recovery replaces assumption-based restoration with continuous verification of data health. Instead of treating backups as inherently safe, organizations evaluate signals across the data lifecycle to determine which recovery points are trustworthy.

“Can we restore?” is not the right question.

Organizations must ask: “Can we restore clean, validated systems under adversarial conditions?”

Modern resilience platforms like Commvault Cloud perform inspection at multiple stages. Before protection, behavioral analytics and threat intelligence help identify suspicious activity in production workloads.

Later, during backup operations, anomaly detection analyzes entropy shifts, unusual file changes, and known threat indicators to help identify potential contamination. Then, after data is stored, continuous scanning helps uncover dormant or delayed threats that may otherwise go unnoticed.

AI-assisted analytics are essential at this scale. They help correlate signals over time, surface high-confidence risks, and reduce manual investigation overhead. Importantly, these capabilities must operate before backup, during backup, and finally during recovery.

This layered approach creates an evidence trail that helps inform recovery decisions with far greater precision.


What Is ResOps for Cyber Recovery?

As cyber recovery becomes more complex and security-sensitive, just stocking up on tools is not enough. Organizations need a repeatable operating discipline that aligns security, IT, and data protection teams around a shared outcome.

This is the foundation of resilience operations (ResOps). ResOps treats recovery as a continuous, measurable operational capability rather than a one-time event. It emphasizes shared intelligence, validated recovery paths, and proven readiness. Some of its key capabilities include helping provide:

  • Continuous visibility through anomaly detection across the data lifecycle.
  • AI-assisted threat detection, threat intelligence, and deception-driven early warning.
  • Restoration of clean data.
  • On-demand, air-gapped environments to validate recovery paths.
  • Repeatable testing and improvement.

A critical element of ResOps is integration with the broader security ecosystem. Modern architectures connect with security information and event management (SIEM); security orchestration, automation, and response (SOAR); extended detection and response (XDR); endpoint; and identity platforms to help enable coordinated response.

SOAR-driven orchestration is especially important during active incidents. Automated playbooks help enable consistent execution, reduce manual errors, and accelerate decision-making across teams.

The adoption of the ResOps model highlights a critical shift in cyber resilience, turning a traditionally siloed process into a highly repeatable engineering capability.


How Does Clean Validation Enable Safe Recovery?

Restoring backups is not as simple as it sounds. Performing restorations hastily can reintroduce malware into production environments. Every restore point must be treated as potentially suspect until proven otherwise.

The necessity for cleanliness of restored data has led to validation gates such as Cleanroom Recovery and Synthetic Recovery.

Cleanroom Recovery helps deliver a fast, on-demand, cloud-based recovery environment for testing, cyber forensics, and recovery staging. This can be practiced by using automated runbooks and pre-configured systems to safely validate workloads before returning them to production.

Complementing this, Air Gap Protect helps provide immutable backups that are stored separately from the production environment, helping prevent attackers from altering or deleting critical recovery data.

AI-assisted Synthetic Recovery extends this validation. It leverages malware and encryption detection to create a curated, composite recovery point that combines the most recent clean version of files across all backups into a single recovery point. This helps reduce the amount of data roll-back or the discarding of good data when performing a recovery.

Together, these mechanisms help transform recovery from a best-effort process into an evidence-backed solution.

Why Is Workload Portability Critical for Cyber Resilience Frameworks?

Enterprise environments now span on-premises infrastructure, multiple public clouds, container platforms, and SaaS ecosystems. Cyber recovery architectures must reflect this reality. Rigid recovery models can create friction and delay, hampering fluid recovery and the safety of backup data.

Any-to-any portability is essential for cyber resilience. 

Any-to-any recovery at enterprise scale means businesses have the flexibility to:

  • Restore workloads across heterogeneous infrastructure.
  • Migrate between cloud providers when needed.
  • Support rebuild-from-scratch scenarios when environments are fully compromised.
  • Support diverse hypervisor migrations and storage platforms.

Portability helps allow recovery decisions to be driven by business priorities rather than platform constraints. It also helps reduce infrastructure lock-in during large-scale incidents.


How Does Minimum Viable Recovery Guide Business Continuity?

When a major cyber incident occurs, attempting to restore everything at once often creates unnecessary delays and complexity. During such scenarios, beginning with an organization’s minimum viable systems and working toward full business recovery can be a powerful strategy.

This approach prioritizes the systems and data required to help restore core business operations first. Recovery sequencing aligns with business impact rather than infrastructure topology. Key elements of this practice include high dependency awareness, tiered recovery objectives, automated runbooks, and continuous testing and refinement.

Minimum viable recovery helps provide a myriad of advantages:

  • Confident, trusted recovery of the most critical parts of the business.
  • Much faster return to continuous business operations.
  • Rapid recovery of identity systems, critical communications applications, and essential data.

Minimum viable recovery helps accelerate time to business continuity. It helps enable organizations to regain operational capability quickly, even if full restoration takes longer. This process also aligns directly with the ResOps philosophy of measurable readiness.

Conclusion: Bringing Together Every Aspect of Cyber Resilience

Present-day cyber resilience is not defined by an organization’s capability to create backups. It is defined by how confidently it can restore trusted operations under real adversarial pressure. It demands an architecture that helps continuously connect detection, validation, isolation, and orchestration.

In a mature architecture, these capabilities reinforce each other in real time. Detection signals help inform Cleanpoint™ confidence. Validation workflows continuously test recoverability. Cleanroom environments help provide a controlled proving ground before production cutover. Orchestrated runbooks help align technical recovery with business priorities. When these elements operate together under a ResOps model, they help provide organizations measurable confidence in their ability to recover.

As cyber threats continue to evolve, the defining advantage will not be how quickly systems can be restored, but how reliably clean, trusted operations can be reestablished at scale.

Frequently Asked Questions

Why are traditional backup strategies no longer sufficient for cyber resilience?

Traditional disaster recovery was designed for outages and hardware failures, not adversarial attacks. Modern ransomware targets backup repositories, corrupts restore points, and disables protection systems. Commvault addresses this by combining Air Gap Protect, anomaly detection, and Cleanroom Recovery to validate and isolate restore points before returning data to production.

What is evidence-driven recovery, and why does it matter?

Evidence-driven recovery uses anomaly detection, threat intelligence, and validation workflows to confirm restore points are clean before deployment. Commvault Cloud implements this through continuous inspection before, during, and after backup — helping surface contamination early and enabling faster, more confident restoration under adversarial conditions.

What is ResOps, and how does it improve cyber recovery?

ResOps is an operating model that treats recovery as a continuous, measurable discipline rather than a one-time event. Commvault supports ResOps by connecting anomaly detection, validated recovery paths, and Cleanroom-based testing into a shared workflow that aligns security, IT, and data protection teams around measurable readiness.

How do Cleanroom Recovery and Synthetic Recovery support safe restoration?

Cleanroom Recovery provides an isolated environment where workloads can be tested and validated before returning to production. Synthetic Recovery uses AI-assisted detection to help assemble the most recent clean versions of files into a verified recovery point. Together, they help prevent reintroducing malware during restoration.

Why is workload portability important during a cyber incident?

In hybrid and multi-cloud environments, organizations need the flexibility to restore workloads across different platforms. Commvault’s any-to-any portability capability allows recovery across heterogeneous infrastructure, cloud migration between providers, and rebuild-from-scratch scenarios — helping ensure recovery decisions are driven by business priorities rather than platform constraints.

What is minimum viable recovery, and how does it support business continuity?

Minimum viable recovery prioritizes restoring the most critical systems required to resume core business operations. Commvault supports this through tiered recovery sequencing aligned to business impact — using automated runbooks and continuous testing to help organizations restore identity systems, critical applications, and essential data before completing a full rebuild.

Explore Related Resources

Commvault’s Complete Cloud Platform

Platform

Cleanroom Recovery

See how Commvault’s on-demand, isolated cloud recovery environment enables safe workload testing, forensics, and production validation after a cyberattack.
Explore the capability about Cleanroom Recovery
IDC MarketScape

A Leader in the IDC MarketScape for Worldwide Cyber-Recovery

Commvault recognized as a Leader for cyber recovery breadth, ecosystem integration, and dedicated cyber-resilience training capabilities.
Read the assessment about A Leader in the IDC MarketScape for Worldwide Cyber-Recovery

Key Takeaways

  • Regulatory pressure, board expectations, and real-world conflict are accelerating the shift from prevention-focused spending to resilience outcomes.
  • The architecture of resilience has grown significantly more complex, especially as AI systems introduce new data lineage and recovery challenges.
  • Cyber resilience must eclipse traditional disaster recovery and treat disruption as a continuous operating condition rather than an exceptional event.
  • Resilience operations (ResOps™) provides an operating model for making resilience continuous, cross-functional, and demonstrable under actual conditions.
  • The hardest part of the transition to ResOps is organizational. Fragmented ownership and misaligned priorities remain the most common failure modes.

Disruptions have become business as usual. Over the past year alone, we’ve seen:

When the operating environment is inherently uncertain, CISOs and CIOs need to rethink their approach to business continuity.

In a recent webinar, David Nowak, principal at Deloitte’s Cyber Risk Service; Kent Meyer, managing director at Deloitte; and Shilpi Handa, IDC’s associate research director for cybersecurity in the META region, joined me to discuss what operational resilience actually demands in strategy, in architecture, and in day-to-day operations.

Sneak Peek: It’s No Longer a Matter of If, but When

In this clip from the webinar, you’ll hear why outages are no longer just IT events – they are business events. Boards and regulators are now shifting focus from if an outage occurs to how quickly organizations can recover.

Why Disaster Recovery Isn’t Enough

Per NIST’s definition, cyber resilience goes beyond traditional security by assuming breaches will happen and focusing on survival and rapid recovery, not just prevention. This assume-breach framing has been part of zero trust for years, but how many organizations are actually putting its implications into practice?

Backup operations focus on whether data has been copied, and disaster recovery on whether systems can be restored, but true resilience demands that you answer a much harder question: Can these services be restored end to end, under stress, and continuously?

We’re seeing this mindset take hold across all sectors, driven by a combination of regulatory pressure and board-level expectations.

  • In Europe, the EU’s Digital Operational Resilience Act (DORA) now mandates specific resilience outcomes and recovery timelines.
  • The North American Electric Reliability Corporation (NERC) Critical Infrastructure Protection regulations are undergoing a resilience lens as well.
  • Securities and Exchange Commission (SEC) breach reporting requirements have put a spotlight not just on disclosure but on what organizations are doing to recover.

Boards now treat any outage as a business harm event, and the expectation has shifted to demonstrating not just that recovery is possible, but that it can happen quickly, with high confidence.

IDC research by Handa illustrates the way traditional recovery operations can fall short. Following the outbreak of war in the Middle East, she found CIOs and CISOs struggling with the continuity not only of technology, but also of people and processes as staff relocate overnight and offices become inaccessible, leaving no one to run manual failover.

And this is just one of the countless unpredictable scenarios organizations need to account for.

Building the Architecture of Resilience

A resilience strategy encompasses both what you protect and whether you can recover it. On the former count, the scope of what needs to be protected has expanded steadily, including identity systems, communication platforms, workloads, productivity tools, CI/CD pipelines, and structured and unstructured data.

The latter point – “whether you can recover it” – creates the new requirements for that strategy. It’s not enough to capture point-in-time snapshots and define traditional recovery objectives. As adversaries target backup infrastructure, organizations must now examine recovered data and confirm that it’s free of compromise before bringing it back online. Isolated recovery environments, air-gap protection for critical services, and cleanroom capabilities have become essential components of resilience architecture.

AI introduces a new layer of difficulty. To recover an AI model, you’ll need not just a backup of the model file itself, but everything that went into creating it. This includes its datasets, hyperparameters, framework versions, feature engineering, and infrastructure configurations, as well as a complete dependency map showing how it all fits together.

Meyer frames cyber resilience solutions as a way to democratize disaster recovery. Whereas traditional disaster recovery was siloed inside IT and accessible only to specialists, newer platforms can give security operations teams, business owners, and operations staff the visibility they need to engage.

That matters for organizations with constrained resources, and it changes what’s possible in terms of moving from tabletop exercises to real, demonstrable restores.

ResOps: The Operating Model for Sustainable Resilience

ResOps treats resilience as a continuous function, not just something that happens in response to an incident.

In simple terms, it’s about operationalizing resilience when normal operating assumptions no longer hold. Instead of taking for granted that your backups will be available, clean, and restorable when disaster strikes, with ResOps you’re continually discovering where data lives, protecting and capturing it across on-premises and cloud environments, detecting anomalies, recovering to a trusted state, and restoring workloads that have been fully validated. That way, you’re increasingly ready for a disruption, and more confident that you’ll be able to get through it successfully.

Nowak offers a phrase that captures the essence of ResOps: resilient by design. It’s the successor to the secure-by-design principle that shaped the last generation of security architecture. Beyond building systems that resist compromise, we’re now building systems that can continue functioning when compromise occurs.

The Human Side of ResOps

Adopting ResOps is at least as much about people and process as it is about technology. As Handa notes, organizations don’t typically fail because they lack the right tools; they fail because ownership is fragmented across too many roles, with no shared operating rhythm and no clear decision rights when things go wrong.

ResOps forces organizations to answer questions many haven’t yet worked through, such as:

  • Who has the authority to restore services if the primary team is unavailable?
  • How should remote execution playbooks be structured?
  • How will cross-regional failover work when it can’t depend on a single location?

In that sense, ResOps is less about restoring systems and more about enabling the continuity of decision-making, execution, and accountability under disruption.

To this end, many organizations have created a chief resilience officer title, particularly in state and local government. Whether filled by the CISO, the CIO, or someone new, the emergence of this role reflects broad accountability beyond IT. It requires an owner with the cross-functional authority and communication skills to bring business leaders, security teams, and operations staff into a shared operating rhythm.

The role also includes translating the case for resilience into terms that resonate across stakeholders, including monetary impact for the board, operational continuity for practitioners, and regulatory compliance for GRC teams. The goal is a decision-rights framework that’s been tested in simulations before it’s needed in an incident.

Putting It All Into Practice

Commvault helps organizations put ResOps into practice, from discovering and protecting data across on-premises and cloud environments, to detecting anomalies, recovering to a clean state, and restoring validated workloads. For organizations working to move from tabletop exercises to demonstrable restores, these are the capabilities that help make resilience operational in uncertain times.

Resources from the session, including IDC research on CIO readiness and Deloitte materials on evidence-based recovery, are available through the on-demand page.

FAQs

Q: What is cyber resilience and how is it different from disaster recovery?

A: Disaster recovery focuses on restoring systems and data after an incident. Cyber resilience is a broader, more active posture: It assumes disruptions will occur and asks whether services can be restored end to end, under stress, on a regular basis. Many organizations have strong disaster recovery plans that nevertheless leave them exposed when a real incident unfolds under unexpected conditions.

Q: What’s driving organizations to prioritize resilience over prevention?

A: Regulatory frameworks like DORA and evolving NERC standards now mandate specific resilience outcomes, not just security controls. As a result, boards are focusing on recovery timelines as a business metric.

Q: What makes AI systems harder to back up and recover than traditional data?

A: Backing up an AI model means capturing more than the model file itself. A model is the product of a specific training process involving datasets, hyperparameters, framework versions, and infrastructure configurations. Without that full context, recovery may produce something that can’t be trusted or reproduced.

Q: What is ResOps, and how does it differ from a traditional resilience program?

A: ResOps is an operating model that treats resilience as a continuous, cross-functional discipline rather than a contingency plan. Where traditional programs tend to be siloed in IT and activated after an incident, ResOps brings together security, operations, business owners, and leadership around shared playbooks, clear decision rights, and ongoing validation of recovery readiness.

Q: What are the biggest obstacles to adopting ResOps at scale?

A: The challenges are organizational, including fragmented ownership, misaligned priorities, and the absence of a shared operating rhythm. ResOps requires agreement – before an incident occurs – between security operations teams, business owners, and operations staff on what’s critical, who’s responsible, and how recovery will be validated.

Michael Thelander is Senior Director, Product Marketing at Commvault.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

Key Takeaways

  • Commvault AirGap is immutable by design, and WORM lock capabilities are available for organizations with additional compliance and regulatory requirements.
  • Claims that AirGap backups are not truly immutable are inaccurate and do not reflect the platform’s documented capabilities.
  • WORM-enabled storage introduces overhead across the industry, but Commvault helps minimize that impact through efficient data management and cloud-native architecture.
  • The cost of backup storage extends beyond capacity consumption and should include infrastructure, compute, and operational expenses.
  • Organizations should help validate backup resilience through real-world testing rather than relying on vendor marketing claims.

You may have recently encountered claims from a competitor suggesting that Commvault AirGap (previously called Commvault Air Gap Protect) contains a critical security gap – that backups are not truly immutable, or that enabling WORM (Write Once, Read Many) lock results in two to three times higher storage costs.

Let us address this directly: These claims are inaccurate.

Immutable and WORM Lock Support in AirGap

AirGap is immutable by design, meaning that once data is written, it cannot be altered – a foundational capability that has been part of the platform since its initial release.

For organizations with regulatory or compliance requirements, Commvault also supports WORM lock capabilities in addition to immutability. These protections are available across  supported cloud storage targets, including:

  • Amazon S3 Object Lock
  • Microsoft Azure Blob immutability policies

These features are documented, rigorously tested, and actively used by customers in production environments today.

In our latest platform release, we further expanded WORM lock support within AirGap. This enhancement extends protection across both cloud and on-premises storage environments, delivering broader and more comprehensive coverage than many competing solutions.

Storage Efficiency: Understanding the Full Picture

Across the industry, one fact remains consistent: WORM-enabled storage introduces some degree of overhead. Because WORM-locked data cannot be modified after it is written, systems have limited ability to optimize or reduce stored data over time. This is not unique to Commvault – it applies universally across vendors.

What differentiates Commvault is how efficiently this challenge is managed. Our platform helps support native cloud immutability (including S3 Object Lock and Azure immutability policies) and maintain an efficient storage overhead.

However, total cost of ownership (TCO) extends beyond storage overhead alone. Architectures that rely on always-on virtual appliances can introduce ongoing compute costs and operational complexity that compound over time.

By contrast, modern, cloud-native approaches prioritize:

  • Efficient data management.
  • Flexible deployment models.
  • Elimination of persistent infrastructure dependencies.

These design principles can result in more predictable, scalable, and sustainable long-term costs.

Presenting this as a choice between insecure backups and excessive storage costs is misleading. It is a false dichotomy that should prompt careful evaluation of vendors making such claims.

A Growing Trend Among Customers

We are seeing a clear pattern: Organizations are increasingly transitioning to Commvault after choosing not to renew with their previous providers. Common reasons may include:

  • Unexpected limitations as environments scale.
  • Performance constraints tied to appliance-based architectures.
  • Rising infrastructure and operational costs.
  • Renewal pricing significantly higher than initial purchase terms.

These challenges are not isolated – they reflect a broader trend in the market. Customers are seeking solutions that offer flexibility, transparency, and long-term value, without hidden trade-offs.

Cutting Through the Noise

In a market often shaped by aggressive claims and unclear comparisons, objective validation can be  critical. That is why we created the Get Real Challenge – a structured, no-cost assessment that enables you to evaluate backup and recovery solutions in a meaningful way.

Through this program, you can:

  • Simulate real-world cyberattack scenarios.
  • Test recovery capabilities using your own data and environment.
  • Evaluate performance without vendor bias or staged demonstrations.

The result is a clear, evidence-based understanding of your organization’s resilience posture. If you want to determine how your backups would perform under real-world conditions, we would be happy to help you get started. Drop us an email at global-sdr@commvault.com.

FAQs

Q: Is Commvault AirGap truly immutable?

A: AirGap is immutable by design, meaning backup data cannot be altered after it is written. This capability has been a core part of the platform since its initial release.

Q: Does AirGap support WORM lock protection?

A: AirGap helps support WORM lock capabilities across supported cloud storage platforms, including Amazon S3 Object Lock and Microsoft Azure Blob immutability policies. Recent enhancements also have helped expand protection across cloud and on-premises environments.

Q: Does enabling WORM lock significantly increase storage costs?

A: WORM-enabled storage creates some level of overhead regardless of the vendor because protected data cannot be modified after it is written. The more important question is how efficiently a platform manages that overhead and its overall impact on long-term costs.

Q: Why is total cost of ownership more important than storage overhead alone?

A: Storage efficiency is only one part of the equation. Organizations also should consider infrastructure requirements, compute costs, operational complexity, and scalability when evaluating the long-term cost of a backup solution.

Q: Why are some organizations moving away from appliance-based backup architectures?

A: As environments grow, organizations often look for solutions that offer greater flexibility, simpler operations, and more predictable costs. Cloud-native approaches can help eliminate dependencies on always-on infrastructure while enabling organizations to scale more efficiently.

Q: How can organizations validate their cyber resilience strategy?

A: The best way is typically through testing. Running realistic recovery exercises and cyberattack simulations enable organizations to understand how their backups will perform under real-world conditions and help identify gaps before an actual incident occurs.

Kash Ansari is Chief Customer Officer Americas at Commvault.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

Over the past few years, I’ve had more conversations about AI than I can count.

Some are focused on potential. Some are focused on risk. Very few are grounded in how AI actually shows up in day-to-day operations.

That’s why I’m writing about an episode of STRIVE I recorded with Ravit Jain, founder and host of “The Ravit Show.”

We didn’t spend time on hype. We didn’t speculate about the future. We focused on what’s happening right now – and what changes when conversational AI is layered on top of unified resilience.

Watch the full episode.

Key Takeaways: What This Shift Really Means

  • Conversational AI helps lower the barrier to cyber intelligence. Leaders can ask complex questions in plain language and get actionable answers.
  • Unified resilience helps reduce fragmentation. Bringing recovery, security, and governance together can change how quickly organizations respond.
  • Trust is the deciding factor in AI adoption. Without transparency and control, AI doesn’t move beyond experimentation.
  • Clean data matters more than speed alone. Recovery isn’t just about getting systems back – it’s about getting back to a trusted state.
  • AI doesn’t replace expertise. It amplifies it by helping to remove friction and accelerate understanding.

From Dashboards to Dialogue

For years, cybersecurity platforms have relied on dashboards, charts, alerts, and reports. And for technical teams, those tools work. But most leaders don’t think in dashboards. They think in questions:

  • Are we exposed?
  • How long will recovery take?
  • What’s the impact if something happens right now?

Ravit and I talked about how conversational AI changes that dynamic. Instead of navigating layers of tooling, teams can interact directly with their environment using natural language. That fundamentally changes who can engage with cyber resilience – and how quickly decisions can be made.

Sneak Peek: Making Cyber Resilience Conversational

Catch a sneak peek of the full episode where we explore how conversational AI shifts cybersecurity from something you interpret to something with which you can directly interact.

Trust Changes Everything

One theme kept coming up throughout our discussion: trust. It’s easy to build an interface that answers questions. It’s much harder to build one that leaders trust in a real incident.

Trust comes down to a few things:

  • Data integrity
  • Access control
  • Transparency
  • Governance
  • Consistency over time

If an executive asks a question about recovery posture, the answer has to be right. It has to be explainable. And it has to be grounded in data that hasn’t been compromised. Without that foundation, conversational AI is interesting, but not operational. With it, it becomes something teams rely on.

Why Unification Matters

Another part of our conversation that stood out was how much complexity still exists in most environments:

  • Different tools for backup.
  • Different systems for security.
  • Different processes for governance.
  • All operating independently.

That fragmentation slows everything down – especially during an incident.

Unified resilience helps change that by bringing those pieces together into a single operational layer. When conversational AI sits on top of that layer, you’re not querying isolated systems anymore. You’re interacting with a connected view of your entire environment. That’s where things can start to move faster and clarity improves. And that’s where recovery decisions become more confident.

This Isn’t About Replacing People

There’s always a question that comes up when AI enters the conversation: What happens to the teams?

Ravit addressed this directly. AI isn’t replacing expertise – it’s extending it. Security teams still define policy. Recovery teams still validate outcomes. And leaders still make decisions.

What changes is how quickly they can get to the information they need – and how clearly they can understand it. Because when you’re in the middle of a cyber event, that clarity matters.

A Shift in How Organizations Operate

There’s also a cultural shift happening when conversational AI becomes part of the workflow:

  • Security discussions can become easier to follow.
  • More stakeholders can participate.
  • Decisions can happen faster.
  • Silos can start to break down.

Instead of cybersecurity being confined to a handful of specialists, it becomes something the broader organization can engage with. To be clear – that doesn’t make it simpler. But it does make it more accessible.

Watch the Full Episode

In the full STRIVE episode, you’ll discover:

  • How conversational AI is actually being used in cybersecurity.
  • What it takes to build trust into AI-enabled systems.
  • Why unified platforms help change recovery outcomes.
  • How organizations can start thinking about this shift.

Watch now.

If you’re thinking about how AI fits into your resilience strategy, it’s worth the time.

FAQs

Q: What is conversational AI in cybersecurity?

A: Conversational AI allows users to interact with security and recovery systems using natural language, making it easier to access insights without navigating complex tools.

Q: How does conversational AI help improve resilience?

A: Conversational AI helps reduce friction in understanding data, can speed up decision-making, and allows more stakeholders to engage in recovery and security discussions.

Q: Why is trust so important for AI adoption?

A: If teams don’t trust the data, the controls, or the outputs, they won’t rely on AI during critical moments.

Q: What does unified resilience mean?

A: Unified resilience refers to bringing together data protection, security, governance, and recovery into a single, integrated approach, rather than managing them separately.

Q: Does conversational AI replace security teams?

A: No. Conversational AI can help teams work more efficiently by making information easier to access and understand.

Q: Where should organizations start?

A: Focus on data integrity, governance, and unifying visibility across systems before layering in conversational capabilities.

Darren Thomson is Vice President and Chief Technology Officer, EMEA, at Commvault

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

Key Takeaways

  • Cyber resilience depends on both technology and the expertise of the professionals responsible for protecting and recovering critical systems.
  • Continuous learning helps partner teams stay current with evolving threats, hybrid environments, and resilience best practices.
  • Commvault Readiverse Academy provides role-based training, hands-on labs, and certifications designed to build real-world cyber resilience skills.
  • Customers increasingly value partners who can help deliver trusted guidance, accelerate recovery readiness, and maximize resilience outcomes.
  • Investing in ongoing education helps strengthen partner capabilities, builds customer trust, and supports long-term business growth.

Organizations are facing increasing pressure to defend against sophisticated and ever-changing threats. But they also need to manage complex hybrid environments, enable rapid recovery, and maintain continuous operations. Technology plays a critical role – but it’s not enough on its own.

True resilience depends on the readiness, experience and ongoing development of the professionals designing, implementing and supporting these environments. As resilience operations continues to mature, trusted expertise has become a critical differentiator.

The Growing Importance of Continuous Learning

Cyber resilience environments are constantly evolving – from expanding hybrid infrastructures to increasingly advanced threats and rising recovery expectations.

For partner professionals across sales, consulting, engineering and support, staying current is no longer optional – it’s essential. Continuous learning helps teams:

  • Strengthen cyber resilience expertise.
  • Build confidence in customer conversations.
  • Stay aligned with evolving technologies and best practices.
  • Prepare for new roles and opportunities.

Whether supporting ransomware readiness, designing recovery strategies or navigating complex data environments, skilled professionals play a critical role in helping maintain resilience and operational continuity.

Building Expertise Across the Partner Ecosystem

At Commvault, we recognize that different partner roles require different skills and learning paths.

That’s why Readiverse Academy was created – offering role-based learning and outcome-focused certifications across sales, engineering, consulting, architecture and support. Beyond just certification, our courses provide practical expertise that drives real customer outcomes.

The Commvault Readiverse Academy delivers:

  • Role-based learning journeys.
  • Self-paced, flexible learning.
  • Hands-on, scenario-based labs.
  • Certifications that validate real-world readiness.

Today, more than 10,000 active learners across the partner ecosystem are building their expertise through Readiverse Academy, reflecting the growing importance of cyber resilience skills across the industry.

Why Expertise Matters to Customers

Customers aren’t just investing in technology – they’re investing in outcomes.

They expect confidence in their ability to protect, manage, and recover systems when disruption occurs. They need trusted experts who can help reduce risk, accelerate recovery, and maximize the value of their investments.

Partners with continuously developing teams are better positioned to help:

  • Accelerate deployments.
  • Align solutions to evolving needs.
  • Improve operational readiness.
  • Strengthen recovery preparedness.
  • Navigate complex resilience challenges.

As cyber resilience becomes mission-critical, expertise becomes a key differentiator.

Investing in the Future of Cyber Resilience

The cyber resilience landscape will continue to evolve – and so will customer expectations.

For partner professionals, continuous learning helps drive long-term growth, credibility, and readiness. For partners, it helps strengthen delivery and build trust. For customers, it helps enable better outcomes.

Explore the Commvault Readiverse Academy, and start building the expertise that sets your team apart – with role-based learning, hands-on training, and certifications designed for real-world impact.

FAQs

Q: Why is technology alone not enough to achieve cyber resilience?
A: Technology provides the tools needed to protect and recover data, but successful cyber resilience also depends on the people using those tools. Skilled professionals are essential for helping design effective strategies, respond to threats, and enable rapid recovery when disruptions occur.

Q: What role does continuous learning play in cyber resilience?
A: Continuous learning helps professionals stay current with evolving cyber threats, changing technologies, and emerging best practices. It also helps build confidence in customer engagements and prepare teams to address increasingly complex resilience challenges.

Q: What is the Commvault Readiverse Academy?
A: Readiverse Academy is Commvault’s learning platform that offers role-based training, hands-on labs, flexible learning paths, and certifications. Its programs are designed to help sales, engineering, consulting, architecture, and support professionals develop practical cyber resilience expertise.

Q: How do customers benefit from working with highly trained partners?
A: Customers gain access to trusted experts who can help reduce risk, improve operational readiness, accelerate deployments, and strengthen recovery preparedness. This expertise helps organizations achieve better outcomes from their cyber resilience investments.

Q: Why are certifications important in the cyber resilience field?
A: Certifications help validate real-world knowledge and readiness, giving both partners and customers confidence in a professional’s capabilities. They also support career development and demonstrate a commitment to maintaining current expertise.

Q: How can partners prepare for the future of cyber resilience?
A: Partners can prepare by investing in ongoing education, developing role-specific expertise, and staying aligned with evolving resilience requirements. Building a culture of continuous learning helps teams remain effective as customer expectations and threat landscapes continue to evolve.

Thomas Kestner is Global Director Partner Solutions & Services Enablement, WW Education Services, at Commvault.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

We announced a strengthened strategic partnership between Commvault and HPE – one grounded in a shared belief that data protection and cyber resilience needed to evolve alongside modern infrastructure.

And if you were in the room for Antonio Neri’s keynote or watched it online, you might remember Commvault being called out on stage.

At the time, it felt like a strong statement of intent.

Today, heading back to Las Vegas, it feels like something more:

Execution. Momentum. And a real opportunity to build modern, resilient IT for customers..

What’s changed in the past year

In the last twelve months, the conversations we’re having with customers have shifted – but so has the environment they’re operating in.

Yes, data is growing. Yes, AI is accelerating. And yes, you absolutely need to have a resilience plan for AI.

But what’s also changed is the nature of the risk.

We’re now entering what many are calling the age of frontier AI – with advanced models like Mythos fundamentally changing how quickly vulnerabilities are discovered and exploited.

You may have seen that in a recent Commvault announcement, we highlighted how these models are compressing what used to be weeks-long exploitation cycles into minutes, dramatically shrinking the window organizations have to respond or recover.

Attacks are becoming more automated, more autonomous, and more immediate.

Which means what you thought you knew might not apply anymore:

  • That you’ll have time to patch before something is exploited
  • That recovery can happen “after the fact”
  • That backup is enough

That’s what’s really changed.

It’s why the conversations we’re having today – with customers, with partners, and across the industry – are less about if something happens and more about how quickly you can recover when it does.

And it’s also why the joint innovation with partners like HPE – bringing to market differentiated new solutions that solve real customer challenges and strengthen our cyber resilience portfolio – is so incredibly valuable.

Three areas where this partnership has evolved

If you step back and look at the past year of this partnership, I’d group our progress with HPE into three clear areas.

#1 – Deeper technical integration where it matters most

We’ve moved well beyond production and protection across storage infrastructure to run-time platforms. So not only do we enable simplified snapshot management and faster recovery across HPE storage technologies like HPE Alletra Storage MP or HPE StoreOnce, we don’t stop at the storage layer. A great example of that is agentless protection for virtual machines (VMs) managed through HPE Morpheus Software.

Virtualization is in a period of real disruption right now. Customers aren’t just evaluating alternatives – they’re actively migrating. And that introduces risk.

What we’ve focused on is helping to make sure protection doesn’t break and can remain consistent during (and after) those transitions.

Agentless protection adds another layer to simplify that – removing dependencies that can slow down or complicate migrations, while helping keep VMs protected across environments.

This level of integration up the stack means customers can accelerate their VM migration strategy confidently and on their own terms, translating to better operational agility, reduced risk, and greater cost-savings.

#2 – Stronger unified go-to-market alignment – and a more complete resilience solution for customers

The second shift has been in how we go to market together – and what we bring to customers as a unified solution stack.

A big part of that is the role of HPE Zerto Software from Commvault.

By integrating HPE Zerto more deeply into Commvault Cloud, we’ve strengthened our platform with continuous data protection and workload resiliency and mobility that enables customers to modernize platforms and rapidly recover workloads to keep their business running after operational disruptions.

And more recently, we introduced a game-changer with Commvault Flex built on HPE infrastructure, a full-stack solution built on:

  • HPE Alletra Storage MP X10000 high-performance, all-flash storage for accelerated recovery for object and file data
  • HPE ProLiant Compute servers for secure, enterprise-grade computing
  • And an industry-leading cyber resilience platform that’s flexible and scalable enough to take advantage of that performance.

Flex solves customers’ challenges in protecting data-intensive workloads like multi-petabyte data lakes that power AI and analytics applications. With Commvault Flex built on HPE technology, customers get an integrated solution that accelerates recovery, simplifies deployment, scales easily, and can help them meet their resilience objectives and recovery SLAs for the foundational data that powers their business.

One more area that’s really come into focus for us over the past year is GreenLake by HPE. As customers push harder into AI, one thing that becomes clear pretty quickly is how infrastructure is delivered and consumed matters just as much as what’s powering it under the hood. There’s a growing need for environments that can scale, adapt, and evolve alongside these AI workloads without adding more complexity. That’s where GreenLake becomes such an important part of the conversation. It’s not just a platform – it’s how many customers are starting to think about building AI-ready infrastructure and become an agentic enterprise. For us, that means doubling down on how Commvault shows up in that ecosystem, continuing to invest in tighter integration and an optimized experience. It’s an area we’re really excited about, and one where you’ll continue to see both teams pushing forward together.

#3 – Real customer outcomes validating the direction

The third area – and probably the most important – is what we’re seeing in customer environments.

We’re starting to see this architecture land in meaningful ways.

For example:

  • A large European bank leveraged the combined Commvault and HPE solution to strengthen cyber resilience across mission-critical banking systems – while also supporting regulatory requirements like DORA compliance. What made the difference here was the combination of Commvault’s architectural advantages and tight integration with the high-performance HPE Alletra Storage MP X10000, enabling the customer to meet recovery objectives that other solutions couldn’t match.
  • A major online gaming organization in South Africa took a slightly different path, focusing on availability and uptime for their platform. In this instance, integrating HPE Zerto into the broader Commvault offering enabled continuous replication and faster recovery, supporting a high-availability environment where even brief disruptions have business impact. The customer got a more complete resilience offering, delivered end-to-end by Commvault for a more streamlined procurement and support process.

Different use cases – but a common theme:

Customers aren’t just buying backup anymore. They’re investing in resilience as part of their production architecture.

Why hybrid infrastructure matters more than ever

If you zoom out, the pattern is clear.

AI workloads are amplifying everything. There’s more data, cycles are faster, and there’s less tolerance for disruption.

And increasingly, the limiting factor isn’t compute – it’s data: How quickly it can be accessed, how efficiently it can be moved, and how fast it can be recovered when something goes wrong.

That’s why platforms like the HPE Alletra Storage MP X10000 are playing a bigger role in these conversations – high performance, scale-out storage that can scale to meet extreme capacity and throughput demands. And when it’s integrated in a solution like Commvault Flex, it creates something that’s increasingly important – a protection and recovery layer that can actually keep up with AI.

Looking ahead to HPE Discover

Heading into this year’s event, there’s a different energy.

A year ago, we were talking about what we could build together.

Now, we’re seeing:

  • Deeper technical integration
  • Clearer go-to-market alignment
  • And real customer outcomes that validate the approach

There’s still a lot of work ahead. But it feels like we’re at one of those points where things start to really take off.

Because the reality is simple:

AI doesn’t wait.

And increasingly, neither can your recovery strategy.

If you’re going to be at HPE Discover 2026, I’d encourage you to stop by and take a look.

Have a conversation with our team at our booth.  Take in a demo. Attend our breakout session. Or setup a meeting with our exec teams for a deeper dive.

I can’t wait to see you there – and to see what all this incredible momentum brings in the coming year.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

There are a lot of conversations happening right now about cyber resilience. Most focus on technology: Detection speed. Recovery architecture. AI-enabled security operations.

All of those things matter. But after sitting down with Dr. Erika Voss, SVP, Global Chief Security & Data Officer at Blue Yonder, and Sam Archey, VP of Trust at Blue Yonder, I kept coming back to something much more fundamental: Trust.

Not trust as a slogan or marketing message, but trust as something operational. Something built deliberately over time and tested in the moments when organizations are under the most pressure.

That distinction can matter because resilience today isn’t just about recovering systems. It’s also about how organizations communicate, how they lead, and how they maintain confidence while uncertainty is still unfolding.

And for a company like Blue Yonder – operating at the center of global supply chains – that challenge becomes even more visible.

Watch the full episode.

Key Takeaways: What Modern Cyber Trust Actually Requires

  • Trust is built through consistency, not perfection. Customers don’t typically expect immediate answers, but they do expect transparency and follow-through.
  • Resilience is operational, not theoretical. Communication, coordination, and decision-making processes can matter as much as technical controls.
  • Strong relationships built before an incident can determine how effectively teams respond during one.
  • Supply chain resilience can raise the stakes because disruptions ripple across interconnected ecosystems.
  • Organizations are increasingly judged not on whether incidents happen – but on how they respond when they do.

Resilience and Trust

One thing became clear very early in this discussion: Erika and Sam don’t think about resilience as a standalone security function. They think about it as a trust function.

Most organizations still separate these ideas:

  • Security handles the technical response.
  • Communications manages messaging.
  • Leadership steps in when escalation is required.

But what Blue Yonder has built is much more integrated than that. Their approach recognizes that customer trust is shaped in real time by operational behavior – not just technical outcomes.

And in a supply chain environment, where countless organizations are interconnected, that operational behavior can become incredibly visible. When something breaks inside that ecosystem, the impact rarely stays isolated. 

The Moment Trust Is Actually Tested

One of the strongest themes throughout the conversation was how quickly trust can be lost – and how intentional organizations must be to preserve it.

Erika put it bluntly: Customers are no longer evaluating whether companies experience incidents. That’s become table stakes in the modern threat landscape. What they are evaluating is something much more specific: Did they hear it from you first?

That distinction changes how organizations should think about incident response.

For years, the instinct during cyber events was often to hold communication until every detail was verified. But the reality today is that silence can create uncertainty faster than almost anything else.

Customers don’t typically expect complete answers in the first hour. They want acknowledgment. They want presence. They want to know that someone is actively working on the problem and willing to communicate transparently while things are still unfolding.

That’s where operational trust is built. And according to Erika and Sam, those first 60 minutes can matter more than most organizations realize.

Sneak Peek: Trust Is Critical in a Crisis

In this moment from the STRIVE conversation, we discuss how the first 60 minutes of response can determine customer confidence, reduce propagation delays, and shape long-term business relationships.

Building Trust Before You Need It

The trust Blue Yonder has built with customers wasn’t created during a single crisis. It was built through repeated interactions over time – through transparency, responsiveness, and operational discipline long before pressure entered the equation.

The same applies internally.

One thing both Erika and Sam emphasize is the importance of relationships between teams before incidents occur. Security, communications, engineering, operations, and leadership need to know how to work together ahead of time. Otherwise, the first real test of collaboration happens during a crisis, which can be the worst possible moment to establish operational alignment.

That’s why they spend so much time focusing on process maturity, stakeholder engagement, and tabletop exercises.

Not because those activities are theoretical. Because they create familiarity.

And familiarity helps reduce friction when pressure rises. 

Why Tabletop Exercises Matter More Than Most Organizations Think

There was a particularly practical section of the conversation around tabletop exercises that I think a lot of organizations need to hear.

Too often, tabletops become compliance activities. Something organizations run once or twice a year to satisfy requirements and move on from. But the way Blue Yonder approaches them is much more operational.

For them, tabletops are rehearsals for coordination.

  • Who makes decisions?
  • How does escalation happen?
  • Which external partners need to be involved?
  • How do legal, communications, and engineering interact?

Those questions become incredibly important during live incidents. And if teams haven’t worked through them ahead of time, response can slow down immediately.

Sam described how teams begin to understand what it may actually feel like to be pulled into an incident under pressure. That experience matters because it helps build muscle memory – not just for technical teams, but for leadership and operational stakeholders as well.

The organizations that recover most effectively are rarely improvising everything in real time. They’ve practiced.

The Human Side of Resilience

What I appreciated most about this conversation was how grounded it was in the human reality of resilience work.

Cyber resilience often gets framed entirely through technology. But people often still determine outcomes.

  • How leaders communicate.
  • How teams collaborate.
  • How organizations behave when information is incomplete.

Those factors help shape customer trust just as much as recovery timelines or technical controls do.

And perhaps the most important lesson from Erika and Sam is that trust isn’t earned during easy moments. It’s earned during uncertainty. During ambiguity. During the moments when organizations have to choose transparency over silence and consistency over perfection.

Watch the Full Episode

In this discussion, you’ll discover:

  • How Blue Yonder operationalizes customer trust.
  • Why consistency can matter more than immediate perfection.
  • The role of communication during cyber incidents.
  • How tabletop exercises strengthen resilience.
  • Why supply chain environments can change the stakes for cyber recovery.

Watch now.

FAQs

Q: Why is trust so important in cyber resilience?

A: Because customers increasingly evaluate organizations based on how they respond during incidents, not simply whether incidents occur.

Q: What does “trust as an operating model” mean?

A: It means trust is continuously reinforced through operational behavior, communication consistency, and transparency – not just during crises.

Q: Why do the first 60 minutes of response matter so much?

A: Early communication can shape customer perception, help reduce uncertainty, and help establish credibility during rapidly evolving situations.

Q: How do tabletop exercises improve resilience?

A: They help teams rehearse coordination, escalation, and communication processes before real incidents occur.

Q: What is the biggest lesson from this discussion?

A: That resilience is typically deeply tied to operational trust, and organizations should build that trust before they need it most.

Q: How can organizations help improve customer trust during incidents?

A: By communicating consistently, prioritizing transparency, and building strong internal coordination long before a crisis begins.

Chris Mierzwa is Senior Director, Portfolio Marketing, at Commvault.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

Key Takeaways

  • Cyber resilience in MEDITECH environments goes beyond backup and recovery; it focuses on maintaining care delivery and operational continuity during disruptions.
  • Healthcare organizations face significant ransomware risk, making rapid, reliable recovery essential for clinical operations.
  • Traditional data protection approaches often fail to address the complex dependencies between clinical systems, applications, and workflows.
  • Effective recovery strategies must coordinate the restoration of interconnected systems to minimize downtime and operational impact.
  • Commvault’s MEDITECH-focused approach helps combine snapshot-based protection, automated recovery workflows, and recovery visibility to strengthen resilience and preparedness.

When ransomware or operational disruption impacts clinical systems, the effects can ripple quickly across the organization, disrupting workflows, slowing staff productivity, and putting timely care delivery at risk. In these moments, the ability to recover quickly and confidently becomes just as important as preventing the disruption in the first place.

That is why Commvault’s approach to protecting MEDITECH environments is centered on recoverability, resilience, and operational readiness, not just data preservation.

Healthcare remains one of the sectors most heavily targeted by ransomware. In a MEDITECH environment, downtime can interrupt medication workflows, delay access to diagnostic information, and force staff into manual workarounds that increase both risk and complexity. In this context, a strategy that looks good on paper is not enough. Health systems need confidence that recovery will perform under real-world pressure.

That is where Commvault can make a meaningful difference.

Why Traditional Data Protection Is Not Enough for MEDITECH

Many organizations still rely on data protection approaches built for general IT environments rather than the operational realities of healthcare. MEDITECH recovery requires an understanding of application interdependencies, restoration order, validation checkpoints, and the need to bring clinical systems back online with minimal disruption.

A successful recovery strategy must account for more than simply restoring data. It must support the coordinated recovery of critical systems, applications, and workflows that clinicians depend on every day. Commvault helps organizations navigate that complexity with resilient architecture, streamlined recovery workflows, and greater visibility into recovery readiness.

How Commvault Helps Protect MEDITECH in Practice

Commvault’s approach to MEDITECH protection is designed to align with the operational realities of these environments. Rather than relying on a one-size-fits-all backup model, the solution helps organizations capture application-consistent protection points for critical MEDITECH workloads while helping minimize disruption to production operations.

This architecture provides healthcare organizations with a practical path to faster, more confident recovery. Snapshot-based protection can support rapid restoration for operational resilience, while longer-term backup retention strengthens options for audit, compliance, and broader cyber preparedness. The result is a model that supports both day-to-day recoverability and resilience planning for more severe disruption.

What Cyber Resilience Can Look Like in Practice

When Ransomware Strikes Overnight

Consider a regional hospital facing encryption activity overnight. In that scenario, the ability to recover from immutable backup copies and execute a structured restoration plan can be the difference between prolonged downtime and a controlled recovery. Commvault helps organizations reduce that risk with secure recovery options designed to restore critical systems quickly and cleanly.

When Backup Infrastructure Is Targeted

Attackers increasingly attempt to compromise backup infrastructure before launching ransomware. That makes architectural resilience essential. With immutable protection and isolated recovery options, Commvault helps enable clean recovery points to remain available even when adversaries gain access to production systems.

When Proof of Recovery Is Required

Cyber insurers, auditors, and compliance stakeholders increasingly want proof that recovery capabilities are tested, documented, and operationally sound. Commvault supports that readiness with validation workflows, reporting, and evidence that can help healthcare organizations demonstrate resilience before an incident occurs.

How Commvault Supports MEDITECH Resilience

Commvault helps organizations protect MEDITECH database volumes with application-consistent recovery points, providing a stronger foundation for restoration when clinical systems are impacted.

Recovery Designed Around MEDITECH Dependencies

MEDITECH recovery often involves complex relationships between systems and databases. Commvault’s approach helps support coordinated protection of critical workloads and a recovery model designed to help bring systems back online in the appropriate sequence.

Validated Recovery with Flexible Retention

By combining rapid recovery options with longer-term backup retention, Commvault helps healthcare teams strengthen resilience beyond the initial snapshot window and build a more complete strategy for recovery testing, validation, and preparedness.

Greater Visibility into Recovery Readiness

A strong MEDITECH resilience strategy depends on operational clarity. Commvault helps teams centralize protection workflows, improve visibility into recovery readiness, and simplify the management of critical data protection tasks.

Regulatory and Insurance Readiness

From documented recovery workflows to retention strategies that support audit and compliance discussions, Commvault helps healthcare organizations strengthen compliance documentation and demonstrate a more mature resilience posture.

Why Implementation Matters

Successful resilience in a MEDITECH environment depends on more than selecting the right platform. It also requires alignment with validated deployment requirements, infrastructure compatibility, and a protection design that reflects how MEDITECH systems operate in the real world.

For healthcare organizations, that implementation discipline can be just as important as the recovery technology itself. A well-designed resilience strategy helps enable teams to execute recovery processes as expected when they are needed most.

Why Now

Ransomware threats continue to evolve, and attackers increasingly target backup infrastructure before deploying encryption. At the same time, cyber insurers and compliance stakeholders are asking for evidence of tested recovery capabilities, not just installed tools.

For healthcare organizations running MEDITECH, the need to build resilience before an incident occurs has never been more urgent. Investing in recovery readiness today can help organizations better protect operations, accelerate recovery, and reduce the impact of disruption when it matters most.

In healthcare, recovery readiness ultimately comes down to trust: trust that critical data is protected, trust that systems can be restored in the right order, and trust that resilience has been tested before a crisis occurs.

That is the standard Commvault helps organizations pursue in MEDITECH environments, and the foundation for a stronger, more confident approach to cyber resilience.

Organizations can further strengthen that foundation by working with a Commvault Managed Service Provider that specializes in healthcare. In addition to the technology itself, healthcare teams gain access to expertise that can help align protection strategies to MEDITECH requirements, support implementation with greater confidence, and improve ongoing operational readiness.

For healthcare organizations navigating the complexity of MEDITECH, that combination of resilient technology and healthcare-focused expertise can help accelerate preparedness and improve recovery outcomes when they matter most.

Final Thoughts

Recovery in a MEDITECH environment is about more than restoring systems. It is about restoring the clinical workflows that caregivers rely on to deliver patient care.

As ransomware threats continue to evolve and healthcare organizations face increasing pressure to demonstrate operational resilience, recovery readiness can no longer be treated as a compliance exercise or a backup strategy alone. Organizations that prioritize recoverability, validation, and cyber resilience before an incident occurs are better positioned to reduce disruption, protect patient care, and recover with confidence when it matters most.

Learn how Commvault helps healthcare organizations strengthen MEDITECH resilience, accelerate recovery, and build greater confidence in their ability to withstand cyber disruption. Visit our MEDITECH documentation site for more information.

FAQs

Q: Why is cyber resilience especially important for MEDITECH environments?

A: MEDITECH environments support critical clinical and operational workflows that directly impact patient care. Cyber resilience helps healthcare organizations recover quickly from disruptions while maintaining essential services and minimizing care delivery interruptions.

Q: How does cyber resilience differ from traditional backup and recovery?

A: Traditional backup and recovery focus primarily on restoring data after an incident. Cyber resilience expands that focus to include operational continuity, rapid recovery, and proactive measures that help reduce the impact of disruptions.

Q: Why are traditional data protection solutions often insufficient for healthcare organizations?

A: Many traditional solutions are designed for general IT environments and may not account for the complex interdependencies between healthcare applications, systems, and workflows. As a result, recovery can be slower and more disruptive.

Q: What challenges do ransomware attacks create for healthcare providers?

A: Ransomware can disrupt access to clinical information, delay care delivery, and increase operational complexity. Healthcare organizations need recovery solutions that enable fast restoration and confidence in recovery outcomes.

Q: How does Commvault support MEDITECH protection and recovery?

A: Commvault’s approach is designed around healthcare operational requirements, providing application-consistent protection, automated recovery processes, and visibility into recovery readiness to help reduce downtime.

Q: What benefits do snapshot-based protection and automated recovery provide?

A: Snapshot-based protection can support faster restoration of critical systems, while automated recovery workflows help streamline recovery efforts. Together, they improve operational resilience and strengthen preparedness for future disruptions.

Chris DiRado is Principal, Product Experience, at Commvault.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

When we talk about cyber resilience, the conversation usually centers on technology: tools, platforms, automation. All of that matters. But when something goes wrong, those aren’t the things that determine how well an organization responds.

People are.

In this episode of STRIVE, I sat down with Dr. Jessica Barker, co-CEO of Cygenta and a leading expert on the human and psychological aspects of cybersecurity. I asked her to discuss a part of resilience that doesn’t always get enough attention – the human side.

What happens when pressure rises, when decisions have to be made quickly, and when teams are forced to work together in ways they may not be used to?

Watch the full episode to find out what she had to say.

Key Takeaways: What the Human Side Reveals

  • Technology doesn’t fail alone – people and processes are always part of the outcome.
  • Confidence under pressure comes from preparation, not instinct.
  • Clear decision ownership helps reduce hesitation during incidents.
  • Trust between teams helps accelerate response and recovery.
  • Culture plays a measurable role in resilience – it’s not just tools or architecture.

When the Plan Meets Reality

Every organization has a plan. It’s documented, reviewed, and often approved at the highest levels. But the real test isn’t how that plan reads – it’s how it holds up when people are under pressure.

Because that’s when things change. Decisions don’t always follow the script. Communication isn’t always clean. Priorities shift in real time. And in those moments, resilience becomes less about process and more about behavior.

Sneak Peek: Cyber Resilience as a Cultural Norm

In this moment from the episode, Dr. Barker highlights the importance of aligning cybersecurity with organizational values. Rather than positioning security as a blocker, resilient organizations embed it into culture – as an enabler of productivity, positivity, and business growth.

The Role of Confidence

One of the most consistent themes in this conversation is confidence. Not confidence in the tools – confidence in the people using them.

Teams that perform well during incidents aren’t guessing. They’ve seen similar scenarios before. They’ve practiced. They understand how to respond, even when conditions aren’t ideal.

That confidence shows up in small ways with potentially faster decisions, clearer communication, and less second-guessing. And over time, those small differences can add up to a significantly stronger response. 

Decision-Making Under Pressure

When something goes wrong, speed matters – but clarity matters more.

  • Who can make decisions?
  • What authority do they have?
  • When should they escalate?

If those answers aren’t clear, teams hesitate. And hesitation creates gaps. One of the most important parts of resilience isn’t just defining processes – it’s defining decision ownership. When people know where they stand, they tend to act faster and with more confidence.

Trust Is the Multiplier

Technology can help enable better and faster responses, but trust can accelerate them. In most organizations, teams operate in their own lanes. Security focuses on threats, infrastructure focuses on systems, and operations focuses on recovery.

That separation works – until an incident forces everyone together. That’s where trust becomes critical. Teams that trust each other:

  • Share information more freely.
  • Collaborate more effectively.
  • Focus on outcomes instead of ownership.

Without that trust, even well-designed processes can break down.

Why Preparation Still Matters

It’s easy to assume that strong individuals can carry a response. But even experienced teams rely on preparation, such as tabletop exercises, simulations, cross-team drills, and so on. They’re what build the muscle memory that teams rely on when real incidents occur. Without that preparation, even the most capable teams are forced to improvise.

Watch the Full Episode

In this episode of STRIVE, we explore:

  • How human behavior impacts incident response.
  • Why decision clarity matters under pressure.
  • What separates confident teams from reactive ones.
  • How culture influences recovery outcomes.
  • Where organizations should focus to strengthen resilience.

Watch now.

If you’re thinking about resilience beyond technology, this is a conversation worth your time.

FAQs

Q: Why is the human side of resilience important?

A: Because technology alone doesn’t determine outcomes – people do. Their decisions, communications, and executions under pressure are the keys to success.

Q: What role does preparation play in resilience?

A: Preparation helps build confidence and muscle memory, allowing teams to respond more effectively in real scenarios.

Q: How does trust impact incident response?

A: Trust helps enable faster collaboration, clearer communication, and more efficient decision-making across teams.

Q: Why is decision ownership critical?

A: Without clear ownership, teams hesitate, which can slow response and increase risk.

Q: Can strong tools compensate for weak processes?

A: No. Tools support resilience, but without strong processes and alignment, they can’t deliver effective outcomes.

Q: Where should organizations start improving?

A: Focus on cross-team alignment, clear decision-making structures, and regular scenario-based testing.

Darren Thomson is Vice President and Chief Technology Officer, EMEA, at Commvault.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

For years, recovery planning followed a familiar pattern. Build the plan, document the steps, and assume it will work when needed. For a long time, that approach held up. Hardware failures, isolated outages, even natural disasters – these were scenarios organizations could anticipate and plan for with some level of confidence.

But the equation has changed.

In this episode of STRIVE, I sat down with Commvault’s Jason Cray, Principal Product Experience, to explore a reality we continue to see across organizations of all sizes: Most don’t fail because they lack a recovery plan. They fail because they’ve never proven that plan will hold up under real pressure.

Watch the full episode.

Key Takeaways: Why Recovery Plans Break Down

  • A documented plan isn’t the same as a proven one. If it hasn’t been tested in realistic conditions, it’s still an assumption.
  • Recovery is a team sport. Security, infrastructure, and operations must align – or recovery slows down.
  • Most investment still happens “left of boom.” Prevention matters, but recovery readiness often gets overlooked.
  • Testing exposes gaps and builds confidence. Without it, organizations default to hope.
  • Resilience is an operational discipline. It requires iteration, communication, and continuous improvement.

The Problem With ‘It Should Work’

On paper, recovery looks straightforward. You define when to recover to, what needs to come back, and where it should be restored. The process appears logical, structured, and manageable.

But as Jason points out, that simplicity rarely survives real-world conditions.

Plans are written in controlled environments, but they’re executed in chaos. When an incident hits, teams aren’t calmly stepping through documentation – they’re reacting, troubleshooting, and trying to align in real time. That’s where the gap emerges. Not between tools and technology, but between expectation and execution.

Sneak Peek: Why Plans Fail Under Pressure

In this moment from the conversation, Jason and I break down why having a plan isn’t enough – and what it actually takes to know a plan will work when it matters.

We’ve Seen This Before

What’s interesting is that this isn’t a new problem; it’s a familiar one, just in a different context.

If you go back to the early days of disaster recovery, organizations followed a similar pattern. Plans existed, but testing was inconsistent at best. Jason shared an example of spending an entire night helping a client pass a disaster recovery test they thought they were ready for. The plan looked solid. The execution told a different story.

Over time, organizations adapted. They tested more frequently, introduced failover exercises and, in some cases even ran production from secondary environments to prove readiness. That shift from assumption to validation is exactly what cyber resilience now requires.

The First Breakdown: Communication

If there’s one issue that consistently surfaces, it’s communication.

In many organizations, responsibilities are clearly defined – security handles prevention, infrastructure manages systems, and operations owns recovery. Individually, each team may be doing exactly what they’re supposed to do.

But recovery doesn’t happen in isolation. It depends on how well those teams work together when something goes wrong.

As Jason describes, too often it becomes a handoff model: “We’ve done our part, now it’s someone else’s turn.” That approach introduces delays, confusion, and ultimately risk. During a cyber event, coordination matters more than ownership.

The ‘Left of Boom’ Problem

Another pattern we continue to see is the imbalance in where organizations focus their efforts.

There’s significant investment in prevention – security tools, detection platforms, and defensive strategies designed to stop an attack before it happens. That investment is necessary, and it plays a critical role.

But far less attention is given to what happens after the event.

The assumption is that if enough effort is spent on prevention, recovery becomes a secondary concern. In reality, the opposite is true. At some point, something gets through. And when it does, recovery becomes the defining factor in how an organization responds.

From Hope to Evidence

This is where the mindset needs to shift.

It’s not about adding more tools or rewriting documentation. It’s about moving from a model based on hope to one grounded in evidence.

Jason highlights a key observation: The organizations that handle disruption well aren’t the ones that avoid incidents – they’re the ones that experience less impact when those incidents occur. They’ve tested their processes. They’ve validated their assumptions. They understand where their gaps are.

Most importantly, they’ve built confidence – not by believing the plan will work, but by proving it.

Start Small, Build Momentum

For many teams, the challenge isn’t understanding the problem – it’s knowing where to begin.

The answer isn’t to overhaul everything at once. It’s to start small and build from there.

Focus on one or two critical services. Understand what’s required to recover them. Bring together the teams responsible for those systems and test the process end-to-end. From there, expand the scope and continue refining.

This approach does more than improve recovery – it builds alignment, reinforces communication, and creates the foundation for broader resilience.

The Reality: No Plan Survives First Contact

One of the most honest moments in our discussion was this: Even the best plan won’t work exactly as written.

That’s not a failure – it’s expected.

Jason puts it simply: If you don’t have a plan, you will fail. But even if you do have one, it won’t unfold perfectly in the moment.

What matters is how prepared to adapt your teams are. Testing creates that adaptability. It builds the muscle memory needed to respond effectively when conditions don’t match expectations.

Watch the Full Episode

There’s much more we cover in this STRIVE conversation, including:

  • Why recovery plans often fail despite being well-documented.
  • What differentiates organizations that recover effectively.
  • How communication gaps impact execution.
  • Where to start when improving recovery readiness.
  • Why testing is the foundation of resilience.

Watch now.

If you’ve ever questioned whether your recovery plan would actually work, this is a conversation worth your time.

FAQs

Q: Why isn’t having a recovery plan enough?

A: Because most plans are never validated under real-world conditions. Without testing, they remain assumptions rather than proven strategies.

Q: What causes recovery plans to fail?

A: The most common issues in recovery plans are lack of testing, poor cross-team communication, and gaps between documented processes and real execution.

Q: What does “left of boom” mean?

A: Left of boom refers to the focus on preventing incidents before they occur. Many organizations invest heavily here but underinvest in recovery capabilities.

Q: How often should recovery plans be tested?

A: Recovery plans should be tested regularly and under varied conditions. Testing should simulate realistic scenarios, not just controlled exercises.

Q: Where should organizations start?

A: Start with a small set of critical services, align the responsible teams, and test recovery end-to-end before expanding.

Q: What is the key mindset shift?

A: Moving from hope-based planning to evidence-based validation.

Chris Mierzwa is Senior Director, Portfolio Marketing, at Commvault.

More related posts


Thumbnail_Blog-Testing-Once-a-Year-2026

Testing Once a Year Is Not a Resilience Strategy

Read more about Testing Once a Year Is Not a Resilience Strategy
Thumbnail_Blog-IDC-Resops-2026

From Recovery to ResOps™: Building Enterprise Resilience That Scales

Read more about From Recovery to ResOps™: Building Enterprise Resilience That Scales
Readiverse-Featured-Image-888-x-500

Ready Is Good. Resilient Is Better.

Read more about Ready Is Good. Resilient Is Better.

We’ve spent years focusing on identity security in the context of people – who have access, what they can do, and how to control it. That model made sense when most activity in the environment was driven by human users.

But that’s no longer the case.

Machine identities – applications, services, APIs, and automated workloads – now play a central role in how modern systems operate. They authenticate, communicate, and execute tasks, often without direct oversight. And in many environments, they already outnumber human identities by a wide margin.

In this episode of STRIVE, I sit down with Dan Conrad, Principal Technologist and a fellow Field CTO at Commvault. We take a closer look at what that shift means – not just from a security perspective, but also from a governance standpoint. And we explore why so many organizations are still treating this as a secondary concern.

Watch the full episode.

Key Takeaways: Where the Risk Is Shifting

  • Machine identities are growing faster than human identities, often by orders of magnitude.
  • Governance models haven’t kept pace, creating blind spots in access and control.
  • Visibility is the core challenge. Many teams don’t fully understand how machine identities behave.
  • Privilege sprawl extends beyond users, with machine identities often holding persistent access.
  • Resilience depends on understanding and managing the remit of these machine identities before they become a problem.

The Identity Model Has Changed

For a long time, identity management was relatively straightforward. You could map users to roles, define access policies, and build controls around predictable behavior. Even with complexity, the model was still anchored in human activity.

Machine identities have broken that model.

They’re created dynamically, often as part of development or deployment processes. They interact across systems in ways that aren’t always visible or well-documented or audited. And unlike human users, they don’t follow a clean lifecycle – they aren’t onboarded and offboarded in the same structured way.

That creates a different kind of challenge. It’s not just about controlling access anymore. It’s about understanding how that access is being used, how it evolves, and how it connects across the environment.

Sneak Peek: You Can’t Phish a Non-Human Identity

In this moment from the STRIVE discussion, Dan describes how attackers aren’t targeting non-human identities directly through phishing – they’re malicious actors using compromised human accounts through social engineering, as a steppingstone to escalate privileges and impersonate powerful machine identities. Once inside, techniques like pass-the-hash and overprivileged service accounts allow attackers to move laterally and vertically, even after passwords are reset.

The Governance Gap

The real issue isn’t that machine identities exist , it’s how they’re governed.

In most organizations, there’s a clear process for managing human access:

  • Requests are approved.
  • Permissions are reviewed.
  • Changes are tracked.

There’s a level of discipline that comes from years of focus on user identity. However, machine identities often fall outside of that structure. They’re created quickly to support applications or automation. They’re granted the permissions needed to function, sometimes more than necessary. And over time, those permissions persist. These overprovisioned accesses are rarely audited, reviewed, and more importantly rarely reduced.

That’s where the gap forms.

It becomes difficult to answer basic questions about access. Not because the information doesn’t exist, but because it hasn’t been organized or managed in a way that makes it usable.

Visibility Before Control

When organizations start to address this problem, the instinct is often to tighten controls.

  • Limit permissions
  • Restrict access
  • Apply new policies

But control without visibility doesn’t solve much.

If you don’t understand how identities are being used, the business context of it in terms of  where they connect, what they interact with, and how they move across systems, then any attempt to restrict them becomes reactive and could result in slowing down business operations.

That’s why visibility needs to come first.

Once you can see how machine identities behave, patterns start to emerge. You can begin to understand where access is excessive, where dependencies exist, and where risk is concentrated. From there, governance can become more precise and more effective.

A Different Kind of Privilege Problem

Privilege sprawl isn’t new. Most organizations have spent years trying to manage excessive access among human users.

Machine identities introduce a similar issue, but with a different dynamic. Their access is often embedded into systems. It’s persistent, automated, and rarely questioned once it’s in place. That makes it harder to detect and easier to overlook.

And when something goes wrong, those identities can become a pathway for malicious actors to exploit

Where to Begin

For most organizations, the challenge isn’t awareness, it’s knowing where to start.

The first step isn’t a major transformation. It’s building clarity. Understanding how many machine identities exist. Where they’re being created. What permissions they have. How they’re used. And most importantly, confirming that a human user is mapped to a collection of non-human identities for the purposes of auditability and accountability.

Those questions sound simple, but they’re often difficult to answer. And that’s exactly why they matter. Because once you can answer them, you’re no longer operating in the dark.

Watch the Full Episode

In this installment of STRIVE, we go deeper into how machine identities are changing the way organizations should think about access, governance, and resilience. It’s a practical conversation about what’s happening now – and what needs to change moving forward.

Watch now.

Resource

If you’re interested in learning more about this topic, check out this e-book on non-human identities.

FAQs

Q: What is a machine identity?

A: A machine identity is a non-human identity used by applications, services, or systems to authenticate and interact with other resources.

Q: Why are machine identities becoming a bigger risk?

A: Because they are increasing in number, often have persistent access, and are not always governed as strictly as human users.

Q: How are they different from user identities?

A: They operate continuously, are embedded in automated workflows, and often lack structured lifecycle management.

Q: What is the biggest challenge organizations face in governing non-human identities?

A: Visibility. Many teams don’t have a clear understanding of how many machine identities are created, used, or interconnected.

Q: How does this impact resilience?

A: If compromised, machine identities can enable a malicious actor’s rapid movement across systems, making incidents harder to contain and recover from.

Q: Where should organizations start?

A: By identifying machine identities, understanding their permissions, and building governance practices that match their scale and complexity. And most importantly, confirming that a human user is mapped to a collection of non-human identities for the purposes of auditability and accountability

Vidya Shankaran is Field CTO at Commvault.

More related posts


Thumbnail_Blog_Identity-Resilience-Vishing_2026

Are You Ready for the Industrialized Vishing Attack?

Read more about Are You Ready for the Industrialized Vishing Attack?
Thumbnail_Blog-Identity-Resilience-MachineID-2026-Linkedin

The Machine Identity Blind Spot Is Now a Primary Attack Surface

Read more about The Machine Identity Blind Spot Is Now a Primary Attack Surface
Thumbnail_Blog-Help-Desk-2026-Linkedin

When the Help Desk Becomes the Front Door to Your Entire Network

Read more about When the Help Desk Becomes the Front Door to Your Entire Network

For decades, IT operations focused on uptime:

  • Keep the infrastructure running.
  • Hit your recovery time objective (RTO).
  • Meet your recovery point objective (RPO).

But modern cyber threats don’t respect infrastructure boundaries – and recovery isn’t just about restoring systems anymore. It’s about restoring clean, trusted data – across teams, under pressure.

In this episode of STRIVE, I sat down with Stephen Foskett, founder and president of the Futurum Group’s Tech Field Day, to discuss an emerging discipline: resilience operations – or ResOps.

And it’s more than a buzzword. It’s a shift in how organizations think about recovery intelligence.

Watch the full episode.

Key Takeaways: What ResOps Changes

  • ResOps moves recovery from infrastructure-focused to business-focused. It’s not just about bringing systems back online – it’s about restoring trusted, usable data.
  • Traditional RTO and RPO metrics aren’t enough anymore. Mean Time to Clean Recovery (MTCR) is emerging as a more meaningful way to measure resilience.
  • Breaking down silos is foundational to cyber readiness. Security, infrastructure, and DevOps must operate in sync – not in parallel.
  • Resilience is an operational discipline, not a tool. Culture, communication, and coordination matter as much as technology.
  • Recovery intelligence is becoming a competitive differentiator. Organizations that recover cleanly and quickly protect revenue, reputation, and trust. 

From IT Ops to ResOps: What’s Changed?

Stephen reflects on an earlier era of IT where teams often supported systems without fully understanding the business applications they powered. Recovery meant restoring infrastructure. Today, that model falls short. Modern environments are:

  • Distributed
  • Cloud-based
  • DevOps-driven
  • Security-sensitive
  • Deeply integrated with revenue streams

ResOps acknowledges that recovery is no longer an isolated IT function. It’s a cross-functional discipline that helps connect infrastructure, software development, and security with real business outcomes.

Why Traditional Metrics Don’t Tell the Whole Story

RTO. RPO. These metrics have guided disaster recovery planning for years. But as Stephen explains, restoring quickly isn’t enough if the data you restore isn’t clean.

Enter a more meaningful metric: MTCR. It’s not just how fast you recover; it’s how fast you can recover to a verified, trusted state.

In a ransomware event, that difference matters enormously. Restoring compromised data can restart an attack cycle. ResOps focuses on restoring operational integrity – not just functionality.

Sneak Peek: Why Clean Recovery Matters

In this moment from STRIVE, Stephen explains why traditional recovery metrics miss the mark – and why recovery is a cross-functional discipline.

The Real Barrier: Organizational Silos

Technology isn’t usually the biggest blocker to resilience. Structure is. Security teams often report to one executive. Infrastructure teams to another. Application teams to yet another. Each with different priorities, different incentives, and different definitions of success.

ResOps challenges that fragmentation.

Stephen discusses how collaborative workshops and cross-functional alignment are helping break down those silos. Because during a cyber event, organizational misalignment slows recovery more than tooling gaps ever will.

Why Commvault Is Leaning Into This Conversation

STRIVE isn’t about product features. It’s about how recovery thinking is evolving. ResOps aligns closely with what we see in the field:

  • Customers struggling with coordination during incidents.
  • Organizations restoring infrastructure but questioning data integrity.
  • Leadership asking for metrics that reflect real business impact.

The concept of MTCR reframes recovery intelligence around business trust – and that’s where the industry is heading. Recovery is no longer a back-office process. It’s an executive concern.

The Future of Recovery Intelligence

Looking ahead, ResOps is likely to mature rapidly. Over the next 12–18 months, organizations are expected to:

  • Integrate security and recovery workflows more tightly.
  • Adopt new recovery-focused metrics.
  • Operationalize resilience earlier in application lifecycles.
  • Invest in intelligence that distinguishes clean data from compromised data.

Cyber threats are accelerating. Recovery strategies must evolve at the same pace. ResOps helps provide a framework for doing that.

Watch the Full Episode

In this episode, we discuss:

  • How ResOps differs from traditional IT operations.
  • Why MTCR is helping to reshape recovery metrics.
  • What organizational alignment looks like in practice.
  • How DevOps culture influences resilience.
  • Where recovery intelligence is expected to head next.

Watch now.

If you’re responsible for cyber readiness, continuity, or recovery strategy, this is a must-watch discussion.

FAQs

Q: What is ResOps?

A: ResOps (Resilience Operations) is an emerging discipline that integrates IT operations, security, DevOps, and business stakeholders to help improve recovery intelligence and organizational resilience.

Q: How is ResOps different from traditional IT operations?

A: Traditional IT ops focuses primarily on infrastructure uptime. ResOps expands that focus to include clean data recovery, cross-functional coordination, and business alignment.

Q: What is Mean Time to Clean Recovery (MTCR)?

A: MTCR measures how quickly an organization can restore verified, clean data and resume safe operations after a cyber event – not just how quickly systems are brought back online.

Q: Why are metrics like RTO and RPO insufficient in modern environments?

A: They measure speed and data currency, but not data integrity. In ransomware scenarios, restoring compromised data can extend disruption.

Q: How can organizations start implementing ResOps?

A: Begin by:

    • Aligning security, infrastructure, and DevOps teams.
    • Evaluating recovery metrics beyond RTO/RPO.
    • Testing clean recovery processes.
    • Breaking down operational silos.
    • Incorporating resilience thinking earlier in system design.

Q: Why is recovery intelligence becoming more important?

A: As cyber threats grow more sophisticated, the ability to recover cleanly, quickly, and confidently directly impacts revenue, customer trust, and regulatory posture.

Darren Thomson is a Field CTO at Commvault.

More related posts


Thumbnail_Blog-Testing-Once-a-Year-2026

Testing Once a Year Is Not a Resilience Strategy

Read more about Testing Once a Year Is Not a Resilience Strategy
Thumbnail_Blog-IDC-Resops-2026

From Recovery to ResOps™: Building Enterprise Resilience That Scales

Read more about From Recovery to ResOps™: Building Enterprise Resilience That Scales
Readiverse-Featured-Image-888-x-500

Ready Is Good. Resilient Is Better.

Read more about Ready Is Good. Resilient Is Better.

Key Takeaways

  • Frontier AI is collapsing vulnerability remediation windows, prevention alone can no longer guarantee security.
  • The question that boards, regulators, and insurers are now asking is not “Do we have backups?” but “Can we prove we can recover cleanly?”
  • Backups are not recovery: a copy tells you data exists, not whether it is clean or restorable.
  • Mean Time to Clean Recovery (MTCR) must become a board-level, continuously measured number – not a theoretical estimate.
  • An Isolated Recovery Environment – air-gapped, immutable, hardened, and identity-isolated – is the baseline, not an advanced capability.
  • What counts as “clean” will keep changing as AI models grow more capable of finding compromises humans cannot anticipate.

I have spent a large part of my career running production systems. I know backup environments from the inside, the ones customers actually trust. I know recovery plans as the things that only reveal their weaknesses when something has already gone wrong. That experience changes how you think about cyber resilience.

From a distance, backup and recovery sounds manageable. Protect the data, store copies, document the runbook, test when you can, restore when you need to. But anyone who has run these environments at scale knows the harder truth: Recovery is where assumptions go to be tested. And right now, too many organizations are operating on assumptions that no longer fit.

For years, security operated on a familiar sequence: Find the vulnerability, patch it, harden the environment, monitor for activity. That model still matters. But the window it depends on is collapsing.

Frontier AI has changed the velocity of vulnerability discovery, attack path chaining, and exploit generation. Models like Claude Mythos and GPT-5.5-Cyber have already demonstrated what this looks like, so far in controlled, early-access testing that still relied on human expertise and carried meaningful false-positive rates, but the trajectory is unmistakable. As access widens, the same capability moves into attackers’ hands.

In a single month, Palo Alto Networks disclosed 26 CVEs, representing 75 underlying issues, after adopting frontier AI models for code scanning, compared with its typical volume of fewer than five CVEs per month.

Researchers also are warning that AI-assisted discovery is collapsing remediation windows, with some exploits now emerging within minutes of disclosure. When the patch window disappears, the remediation math stops working. Prevention cannot carry the full weight of readiness.

Prevention still matters, but it no longer defines readiness. The customers I talk to are not asking whether they need more controls. They already know they do. They are asking whether their business can recover cleanly when those controls fail, when attackers move faster than remediation cycles, or when compromise has been present longer than anyone realized.

That question is now what boards, regulators, and insurers are forcing. They have moved past “Do we have backups?” and toward something more consequential: “Can we prove we can recover cleanly?”

That proof starts with one distinction most organizations still get wrong: Backups are not recovery.

A backup tells you a copy exists. It does not tell you whether the data is clean, whether application dependencies are intact, whether identity services can be safely restored, or whether the recovery sequence still reflects the current environment.

I have reviewed plans that looked complete until someone tried to execute them. The runbook was there, but outdated. The restore worked but took three times longer than the estimate. The system came back, but downstream applications could not connect. None of that is unusual. It is exactly what real testing is supposed to surface. The problem is most organizations discover these gaps during an actual incident.

The metric that matters most when something goes wrong is how quickly you can return to a known-good state. That is why Mean Time to Clean Recovery (MTCR) needs to become a board-level number, not a theoretical estimate in a plan, but a measured, validated time.

The Moving Target: What’s Clean Today May Not Be Clean Tomorrow

With Frontier AI models, the honest answer is this: you cannot guarantee that every vulnerability will be found and remediated in time. Attackers leveraging the same models are discovering and chaining exploits faster than any remediation program can realistically keep pace with. That is not a failure of your security team. It is the new physics of the threat landscape.

What you can control is your ability to recover. That means an Isolated Recovery Environment – backups air-gapped from the internet, unreachable from the production network, and protected from the lateral movement that defines a sophisticated breach. It means immutability and compliance lock, so no credential, however privileged, can shorten retention or delete data outside an authorized process. And it means ResOps in practice: not just backing up data, but continuously testing recovery, automating integrity validation, and measuring your MTCR – the validated time to return to a known-good state.

But here is the part most organizations are not yet accounting for: what counts as “clean” is not a fixed line. As AI models grow more capable, they will increasingly find vulnerabilities that the human mind simply cannot anticipate, novel attack paths, dormant implants, subtle corruptions embedded long before detection. A recovery point that is clean by today’s standards may carry compromise that tomorrow’s AI-assisted forensics will surface. That means your definition of clean must evolve continuously. MTCR is not a number you set once. It is a discipline you maintain, revisiting what clean means, updating your validation criteria, and treating resilience as a living standard rather than a certification you pass once.

So what is a good MTCR? Based on what I have seen work in practice, the target for your entire minimum viable company – the smallest set of systems that lets you keep operating, which I define precisely below, should be under six hours. Six hours is achievable with the right architecture: an IRE ready to run, a pre-validated recovery sequence, and runbooks that are executable rather than readable. If your current MTCR is measured in days, the gap is almost always one of those three.

Four Steps to Stay Resilient in the Frontier AI Era

Accepting that prevention alone is not enough is the starting point. From there, the work gets specific. Here is where I tell organizations to focus.

1. Evaluate your actual recovery risks.

Most recovery risk assessments ask the wrong questions. “Do backups exist?” is not the same as “Can we recover cleanly?” The harder questions are: Can critical systems be restored without reintroducing the threat? Are recovery environments isolated from compromised production systems? Are recovery plans mapped to current dependencies – not the architecture from two years ago?

In a fast-moving vulnerability environment, the gap between “we have backups” and “we can recover” is where organizations get hurt. Assessing that gap honestly, before an incident forces the issue, is where resilience planning must start. That assessment needs to include a business impact analysis: which systems have a recovery window measured in minutes, which in hours, and which can wait a day. Without that tiering, every system looks equally urgent during an incident, and nothing gets restored fast enough.

2. Make isolated recovery and air gapping the baseline – not the exception.

If you are still treating air-gapped, immutable copies as an advanced capability rather than a standard requirement, that assumption no longer holds. When exploitation timelines compress to minutes, you need fallback options that are structurally separated from production identity, network, and management planes – logically or physically isolated, immutable, and with no live path back to production that an attacker can follow.

The goal is not just protection from the current threat, but maintaining clean recovery options when a vulnerability you have not patched yet gets exploited. That happens now. Plan for it.

Isolation only holds if the infrastructure around it is hardened. That means backup infrastructure on hardened operating systems, not generic images, and ideally on physical servers that survive a hypervisor-layer attack. It means encryption keys stored outside the backup platform, in an external vault with just-in-time access and no dependency on production Active Directory. And it means treating your backup domain as a separate identity boundary: no trust to production AD, mandatory MFA, and multi-person authorization for destructive operations. None of this is exotic, it is the baseline for your environment to recover into an uncompromised space.

Equally important is the question of what you are recovering from. Industry incident-response data consistently puts median breach dwell time in the range of weeks, not days. That means your recovery copies need to reach back far enough to find a genuinely clean point, not just yesterday’s backup. Critical systems warrant multiple geographically separated copies, including at least one immutable copy and one that is fully offline. Retention policy is not a storage cost decision. It is a security decision.

3. Know which systems the business cannot operate without – and recover those first.

Most organizations discover their recovery sequence during an incident. That’s why the first 24–48 hours aren’t spent restoring systems, they’re spent deciding what matters.

Organizations know they have to recover identity platforms, billing systems, operational databases, and core infrastructure. What they often have not mapped is the order, the dependencies between those systems, and the downstream applications that cannot function until specific services are back.

This gets more complex as AI becomes embedded in business operations. Data pipelines, model repositories, vector databases, agentic workflows – these are now operational dependencies, not just technical infrastructure. If your recovery sequencing does not account for them, your recovery time estimates are probably wrong.

Defining what it means to operate as a minimum viable company (the smallest set of systems required to keep the business running) and building recovery around that definition is not a theoretical exercise. It is the practical answer to the question every executive team will ask during an incident: What do we bring back first?

In my experience helping customers through active incidents, the first 12 hours answer that question whether you have planned for it or not – what gets recovered in that window becomes your MVC by default. The organizations that come through fastest decided in advance: they knew exactly which systems had to be back within 12 hours and had validated they could do it. If your MVC does not fit in 12 hours, it is not your MVC, it is a wish list. The work is to keep trimming until what remains can realistically be restored in that window, then test it until you can prove it.

4. Automate resilience and test continuously – not on a calendar schedule.

A recovery plan that lives in a document and is reviewed annually is not a recovery capability. It is a hypothesis that has never been tested against reality.

The problem with calendar-based testing is what it misses between cycles. Environments change constantly: new workloads, updated dependencies, infrastructure that has drifted from what the runbook describes. By the time the annual test runs, it is validating a snapshot of an environment that no longer exists. In a threat landscape where exploitation can happen within minutes of disclosure, that lag is not acceptable.

Threat scanning, clean recovery point identification, dependency-aware restoration, and recovery orchestration all need to be automated and running continuously. Not because automation is a best practice, but because the manual alternative cannot keep pace with how fast things now move.

Continuous testing also depends on continuous detection. You cannot select a clean recovery point if you do not know when the compromise began. That is why threat detection, anomaly scanning of backup data, and recovery-point analysis have to feed each other: detection tells you which copies predate the intrusion, and that determination drives which point you actually recover from. Without that link, you are restoring to a date you hope is clean rather than one you have verified, and in a Frontier AI threat landscape, hope is not a recovery strategy.

What continuous testing surfaces is different from what annual tests find. Calendar tests tend to confirm the plan works under controlled conditions. Continuous testing finds the dependency that changed last month, the recovery sequence that breaks when a specific workload is added, the identity service that takes twice as long to restore as the estimate assumed.

Those are the gaps that matter during a real event, and the only way to find them before an incident does is to be testing all the time.

Testing also needs to happen in the right environment. A recovery test that runs against production infrastructure does not tell you whether you can recover when production is compromised. Cleanroom testing – validating restoration in a fully isolated environment with no connectivity back to production – is how you confirm your backup copies are genuinely usable under incident conditions. That includes recovering identity services, external key management, and Tier 0 applications in isolation, with dedicated break-glass accounts that exist outside your normal directory.

What makes daily testing viable is validate restore, a recovery type that exercises the full restore path for every critical asset without touching production. Your backup platform needs to support this natively; if it cannot run an automated, non-disruptive recoverability test across your MVC every day, you do not actually know whether your backups work. In Commvault, this restores against your critical-asset groups, with automated reporting on the recovery status of every protected system.

The same applies to your runbooks. A runbook that lives in a Word document or PDF is a reference manual, not an operational tool – it assumes someone has the time, clarity, and access to read it under pressure. Real runbooks are digital scripts that execute the recovery sequence and validate each step, confirming the application actually works before moving on: not the service started” but “the application responded correctly to a synthetic transaction.” Commvault’s Cleanroom Runbooks are built for this – executable workflows that drive an end-to-end recovery in an isolated environment without a human interpreting a document at every step.

One final point that rarely makes it into recovery plans until it is too late: during a serious incident, your corporate communications infrastructure may itself be compromised or unavailable. Email, Teams, and Slack run on the same infrastructure attackers target. Know in advance which out-of-band channels your team will use to coordinate, and make sure those channels are tested alongside your technical recovery procedures.

Hear more from Commvault’s Chief Security Officer Bill O’Connell on the four critical steps for resiliency in the AI era.

Resilience Is an Operating Discipline, Not a Project

The organizations that will hold up under frontier AI-accelerated threats are the ones that treat resilience as an operating discipline — measured MTCR, continuous validation, and a recovery capability they have proven, not assumed.

The problem isn’t that attacks are getting faster. It’s that recovery hasn’t caught up, and until it does, the math doesn’t work.


FAQs

Q: What is Mean Time to Clean Recovery (MTCR) and why does it matter?
A: MTCR measures how quickly an organization can return to a verified, known-good state after a cyberattack – not just restore data, but confirm it is clean and that application dependencies are intact. It should be a board-level metric with a measured, validated time, not a theoretical estimate buried in a recovery plan. The target for a well-architected MVC – covering all identity systems, critical applications, and isolated environment readiness – is under six hours.

Q: What is an Isolated Recovery Environment and how is it different from a standard backup?
A: An Isolated Recovery Environment is a fully air-gapped, immutable copy of critical data that is structurally separated from production networks, identity systems, and management planes. A standard backup tells you a copy exists. An IRE tells you that copy is protected from the same attack that hit your production environment.

Q: How do we know whether we can actually recover today?
A: The only honest answer comes from testing, not documentation. If you cannot point to a recent, validated recovery of your minimum viable company – ideally a daily automated test – then you do not know, you are assuming. A defensible answer to the board is a measured MTCR backed by continuous validation, not a recovery plan that looks complete on paper.

Q: What do regulators and cyber insurers now expect?
A: The bar has moved from “Do you have backups?” to “Can you prove you can recover cleanly, and how fast?” Regulators increasingly expect demonstrable recovery capability and tested resilience; insurers increasingly price coverage – and pay claims – based on evidence of isolated, immutable backups and validated recovery times. A measured MTCR and a documented testing cadence are becoming table stakes for both.

Rajiv Kottomtharayil is Chief Product Officer at Commvault.

More related posts


Thumbnail_Blog-Clumio-Chat-2026

Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection

Read more about Meet Clumio Chat: An AI Assistant to Help Evaluate Cloud-Native Data Protection
Thumbnail_Blog-Clumio-Fedramp-2026

Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone

Read more about Clumio Advances Cloud-Native Cyber Resilience with FedRAMP® Milestone
Thumbnail_Blog_Agentic-Ransomware-Attack

Cyber Resiliency for AI and Ransomware Recovery

Read more about Cyber Resiliency for AI and Ransomware Recovery

Blog

Protecting AI Workloads: How Can Organizations Achieve Resilience in the AI Era?

AI resilience helps enable protection, recovery, and governance of AI workloads, data, and models by combining threat detection, clean recovery, and controlled data access.

Frequently Asked Questions

What is AI resilience?

AI resilience is the ability to protect, recover, and govern AI systems across their full lifecycle. Commvault’s Protect and Leverage AI capabilities help verify that data, models, and pipelines remain secure, recoverable, and trustworthy — even when disrupted by cyber threats, failures, or operational complexity in hybrid and multi-cloud environments.

Why is protecting AI workloads important?

AI workloads rely on distributed data, models, and infrastructure, making them vulnerable to threats like data poisoning and model corruption. Protecting them helps uphold data integrity, reduce operational risk, and maintain trust in AI-enabled business processes. Commvault helps address these challenges with Metallic AI, unifying ML-driven detection, guided recovery, and automation across Commvault Cloud.

What does full AI stack protection include?

Full AI stack protection safeguards data pipelines, vector databases, models, metadata, configurations, and compute infrastructure. Commvault Cloud Unity covers this breadth — including unified data platforms like Amazon Redshift and Google BigQuery, vector retrieval systems, and compute infrastructure — enabling complete and consistent recovery of AI workloads across hybrid and multi-cloud environments.

Why does clean recovery matter in AI environments?

Clean recovery confirms that restored data is free from corruption, malware, or inconsistencies. In AI systems, compromised data leads to inaccurate outputs and biased decisions. Commvault Synthetic Recovery addresses this by analyzing multiple backup versions to assemble a validated recovery point — so restored AI workloads produce trusted, accurate outputs.

How does AI improve data protection and operations?

Commvault embeds AI across the protection lifecycle — automating threat detection, optimizing backup scheduling, and predicting storage needs through ML-enabled capabilities. Arlie, Commvault’s AI assistant, improves user experience through natural language interactions, guided workflows, and intelligent insights, helping security and IT teams manage complex AI environments more efficiently.

What is responsible AI in data protection?

Responsible AI enables systems to operate with transparency, governance, and control. Commvault supports this through Data Activate — a governed workspace that applies encryption, immutability, and role-based access controls to curate and extend trusted data to AI and analytics platforms, helping prevent misuse and maintain compliance while enabling innovation.