Skip to content

Wirhaben eine verstärkte strategische Partnerschaft zwischen Commvault und HPE bekannt gegeben – – one grounded in a shared belief that data protection and cyber resilience needed to evolve alongside modern infrastructure.

And if you were in the room for Antonio Neri’s keynote or watched it online, you might remember Commvault being called out on stage.

At the time, it felt like a strong statement of intent.

Today, heading back to Las Vegas, it feels like something more:

Execution. Momentum. And a real opportunity to build modern, resilient IT for customers..

What’s changed in the past year

In the last twelve months, the conversations we’re having with customers have shifted – but so has the environment they’re operating in.

Yes, data is growing. Yes, AI is accelerating. And yes, you absolutely need to have a resilience plan for AI.

But what’s also changed is the nature of the risk.

We’re now entering what many are calling the age of frontier AI – with advanced models like Mythos fundamentally changing how quickly vulnerabilities are discovered and exploited.

You may have seen that in einer kürzlich veröffentlichten Mitteilung von Commvault, we highlighted how these models are compressing what used to be weeks-long exploitation cycles into minutes, dramatically shrinking the window organizations have to respond or recover.

Attacks are becoming more automated, more autonomous, and more immediate.

Which means what you thought you knew might not apply anymore:

  • That you’ll have time to patch before something is exploited
  • That recovery can happen “after the fact”
  • Diese Sicherung reicht aus

That’s what’s really changed.

It’s why the conversations we’re having today – with customers, with partners, and across the industry – are less about if something happens and more about how quickly you can recover when it does.

And it’s also why the joint innovation with partners like HPE – bringing to market differentiated new solutions that solve real customer challenges and strengthen our cyber resilience portfolio – is so incredibly valuable.

Drei Bereiche, in denen sich diese Partnerschaft weiterentwickelt hat

If you step back and look at the past year of this partnership, I’d group our progress with HPE into three clear areas.

#1 – Eine tiefgreifendere technische Integration dort, wo es am wichtigsten ist

We’ve moved well beyond production and protection across storage infrastructure to run-time platforms. So not only do we enable simplified snapshot management and faster recovery across HPE storage technologies like HPE Alletra Storage MP or HPE StoreOnce, we don’t stop at the storage layer. A great example of that is agentless protection for virtual machines (VMs) managed through HPE Morpheus Software.

Virtualization is in a period of real disruption right now. Customers aren’t just evaluating alternatives – they’re actively migrating. And that introduces risk.

What we’ve focused on is helping to make sure protection doesn’t break and can remain consistent during (and after) those transitions.

Agentless protection adds another layer to simplify that – removing dependencies that can slow down or complicate migrations, while helping keep VMs protected across environments.

This level of integration up the stack means customers can accelerate their VM migration strategy confidently and on their own terms, translating to better operational agility, reduced risk, and greater cost-savings.

#2 – Eine stärkere, einheitliche Ausrichtung der Markteinführung – und eine umfassendere Resilienzlösung für Kunden

Der zweite Wandel betrifft die Art und Weise, wie wir gemeinsam auf den Markt treten – und was wir den Kunden als einheitliches Lösungsportfolio bieten. Einen großen Teil dazu trägt die Rolle derHPE Zerto Software from Commvault.

By integrating HPE Zerto more deeply into Commvault Cloud, we’ve strengthened our platform with continuous data protection and workload resiliency and mobility that enables customers to modernize platforms and rapidly recover workloads to keep their business running after operational disruptions.

And more recently, we introduced a game-changer with Commvault Flex built on HPE infrastructureeine bahnbrechende Neuerung eingeführt – eine Full-Stack-Lösung, die auf folgenden Komponenten basiert:

  • HPE Alletra Storage MP X10000 – leistungsstarker All-Flash-Speicher für eine beschleunigte Wiederherstellung von Objekt- und Dateidaten
  • HPE ProLiant Compute-Server für sichere Rechenleistung auf Unternehmensniveau
  • And an industry-leading cyber resilience platform that’s flexible and scalable enough to take advantage of that performance.

Flex solves customers’ challenges in protecting data-intensive workloads like multi-petabyte data lakes that power AI and analytics applications. With Commvault Flex built on HPE technology, customers get an integrated solution that accelerates recovery, simplifies deployment, scales easily, and can help them meet their resilience objectives and recovery SLAs for the foundational data that powers their business.

One more area that’s really come into focus for us over the past year is GreenLake by HPE. As customers push harder into AI, one thing that becomes clear pretty quickly is how infrastructure is delivered and consumed matters just as much as what’s powering it under the hood. There’s a growing need for environments that can scale, adapt, and evolve alongside these AI workloads without adding more complexity. That’s where GreenLake becomes such an important part of the conversation. It’s not just a platform – it’s how many customers are starting to think about building AI-ready infrastructure and become an agentic enterprise. For us, that means doubling down on how Commvault shows up in that ecosystem, continuing to invest in tighter integration and an optimized experience. It’s an area we’re really excited about, and one where you’ll continue to see both teams pushing forward together.

#3 – Echte Kundenergebnisse, die die Richtung bestätigen

The third area – and probably the most important – is what we’re seeing in customer environments.

We’re starting to see this architecture land in meaningful ways.

For example:

  • A large European bank leveraged the combined Commvault and HPE solution to strengthen cyber resilience across mission-critical banking systems – while also supporting regulatory requirements like DORA compliance. What made the difference here was the combination of Commvault’s architectural advantages and tight integration with the high-performance HPE Alletra Storage MP X10000, enabling the customer to meet recovery objectives that other solutions couldn’t match.
  • Ein großes Online-Gaming-Unternehmen in Südafrika schlug einen etwas anderen Weg ein und legte den Schwerpunkt auf die Verfügbarkeit und Betriebszeit seiner platform. In diesem Fall ermöglichte die Integration von HPE Zerto in das umfassendere Commvault-Angebot eine kontinuierliche Replikation und eine schnellere Wiederherstellung und unterstützte so eine Hochverfügbarkeitsumgebung, in der selbst kurze Unterbrechungen geschäftliche Auswirkungen haben. Der Kunde erhielt ein umfassenderes Resilienz-Angebot, das von Commvault durchgängig bereitgestellt wurde und so für einen optimierten Beschaffungs- und Supportprozess sorgte.

Different use cases – but a common theme:

Customers aren’t just buying backup anymore. They’re investing in resilience as part of their production architecture.

Warum hybride Infrastruktur heute wichtiger denn je ist

If you zoom out, the pattern is clear.

AI workloads are amplifying everything. There’s more data, cycles are faster, and there’s less tolerance for disruption.

And increasingly, the limiting factor isn’t compute – it’s data: How quickly it can be accessed, how efficiently it can be moved, and how fast it can be recovered when something goes wrong.

That’s why platforms like the HPE Alletra Storage MP X10000 are playing a bigger role in these conversations – high performance, scale-out storage that can scale to meet extreme capacity and throughput demands. And when it’s integrated in a solution like Commvault Flex, it creates something that’s increasingly important – a protection and recovery layer that can actually keep up with AI.

Ein Ausblick auf die HPE Discover

Heading into this year’s event, there’s a different energy.

A year ago, we were talking about what we could build together.

Now, we’re seeing:

  • Eine tiefgreifendere technische Integration
  • Eine klarere Abstimmung der Markteinführung
  • Und konkrete Kundenergebnisse, die den Ansatz bestätigen

There’s still a lot of work ahead. But it feels like we’re at one of those points where things start to really take off.

Because the reality is simple:

AI doesn’t wait.

And increasingly, neither can your recovery strategy.

If you’re going to be at HPE Discover 2026, I’d encourage you to stop by and take a look.

Have a conversation with our team at our booth.  Take in a demo. Attend our breakout session. Or setup a meeting with our exec teams for a deeper dive.

I can’t wait to see you there – and to see what all this incredible momentum brings in the coming year.

More related posts


Thumbnail_Blog-Architect-for-tomorrow-2026

Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience

Read more about Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience
Thumbnail_Blog-Dangerous-Silos-IDC-Resops-2026 (1)

The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan

Read more about The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan
Thumbnail_Blog-IDC-Resops-2026

Business Continuity Planning for the Cloud-Native Era

Read more about Business Continuity Planning for the Cloud-Native Era

There are a lot of conversations happening right now about cyber resilience. Most focus on technology: Detection speed. Recovery architecture. AI-enabled security operations.
All of those things matter. But after sitting down with Dr. Erika Voss, SVP, Global Chief Security & Data Officer at Blue Yonder, and Sam Archey, VP of Trust at Blue Yonder, I kept coming back to something much more fundamental: Trust.
Not trust as a slogan or marketing message, but trust as something operational. Something built deliberately over time and tested in the moments when organizations are under the most pressure.
That distinction can matter because resilience today isn’t just about recovering systems. It’s also about how organizations communicate, how they lead, and how they maintain confidence while uncertainty is still unfolding.
And for a company like Blue Yonder – operating at the center of global supply chains – that challenge becomes even more visible.
Watch the gesamte Folge.

Das Wichtigste auf einen Blick: Was modernes Cyber-Vertrauen tatsächlich erfordert

  • Trust is built through consistency, not perfection. Customers don’t typically expect immediate answers, but they do expect transparency and follow-through.
  • Resilienz ist eine praktische Angelegenheit, keine theoretische. Kommunikation, Koordination und Entscheidungsprozesse können genauso wichtig sein wie technische Kontrollmaßnahmen.
  • Starke Beziehungen, die bereits vor einem Vorfall aufgebaut wurden, können darüber entscheiden, wie effektiv Teams während eines Vorfalls reagieren.
  • Die Widerstandsfähigkeit der Lieferkette kann die Risiken erhöhen, da sich Störungen auf miteinander vernetzte Ökosysteme auswirken.
  • Organizations are increasingly judged not on whether incidents happen – but on how they respond when they do.

Resilienz und Vertrauen

One thing became clear very early in this discussion: Erika and Sam don’t think about resilience as a standalone security function. They think about it as a trust function.
Most organizations still separate these ideas:

  • Der Sicherheitsdienst kümmert sich um die technischen Maßnahmen.
  • Die Abteilung „Kommunikation“ ist für den Nachrichtenverkehr zuständig.
  • Die Führung greift ein, wenn eine Eskalation erforderlich ist.

But what Blue Yonder has built is much more integrated than that. Their approach recognizes that customer trust is shaped in real time by operational behavior – not just technical outcomes.
And in a supply chain environment, where countless organizations are interconnected, that operational behavior can become incredibly visible. When something breaks inside that ecosystem, the impact rarely stays isolated. 

Der Moment, in dem das Vertrauen wirklich auf die Probe gestellt wird

One of the strongest themes throughout the conversation was how quickly trust can be lost – and how intentional organizations must be to preserve it.
Erika put it bluntly: Customers are no longer evaluating whether companies experience incidents. That’s become table stakes in the modern threat landscape. What they are evaluating is something much more specific: Did they hear it from you first?

That distinction changes how organizations should think about incident response.
For years, the instinct during cyber events was often to hold communication until every detail was verified. But the reality today is that silence can create uncertainty faster than almost anything else.
Customers don’t typically expect complete answers in the first hour. They want acknowledgment. They want presence. They want to know that someone is actively working on the problem and willing to communicate transparently while things are still unfolding.
That’s where operational trust is built. And according to Erika and Sam, those first 60 minutes can matter more than most organizations realize.

Vorschau: Vertrauen ist in einer Krise entscheidend

In this moment from the STRIVE conversation, we discuss how the first 60 minutes of response can determine customer confidence, reduce propagation delays, and shape long-term business relationships.

Vertrauen aufbauen, bevor man es braucht

The trust Blue Yonder has built with customers wasn’t created during a single crisis. It was built through repeated interactions over time – through transparency, responsiveness, and operational discipline long before pressure entered the equation.
The same applies internally.
One thing both Erika and Sam emphasize is the importance of relationships between teams before incidents occur. Security, communications, engineering, operations, and leadership need to know how to work together ahead of time. Otherwise, the first real test of collaboration happens during a crisis, which can be the worst possible moment to establish operational alignment.
That’s why they spend so much time focusing on process maturity, stakeholder engagement, and tabletop exercises.
Not because those activities are theoretical. Because they create familiarity.
And familiarity helps reduce friction when pressure rises. 

Warum Tabletop-Übungen wichtiger sind, als die meisten Organisationen glauben

There was a particularly practical section of the conversation around tabletop exercises that I think a lot of organizations need to hear.
Too often, tabletops become compliance activities. Something organizations run once or twice a year to satisfy requirements and move on from. But the way Blue Yonder approaches them is much more operational.
For them, tabletops are rehearsals for coordination.

  • Wer trifft die Entscheidungen?
  • Wie läuft eine Eskalation ab?
  • Welche externen Partner müssen einbezogen werden?
  • Wie wirken die Bereiche Recht, Kommunikation und Technik zusammen?

Those questions become incredibly important during live incidents. And if teams haven’t worked through them ahead of time, response can slow down immediately.
Sam described how teams begin to understand what it may actually feel like to be pulled into an incident under pressure. That experience matters because it helps build muscle memory – not just for technical teams, but for leadership and operational stakeholders as well.
The organizations that recover most effectively are rarely improvising everything in real time. They’ve practiced.

Die menschliche Seite der Resilienz

What I appreciated most about this conversation was how grounded it was in the human reality of resilience work.
Cyber resilience often gets framed entirely through technology. But people often still determine outcomes.

  • Wie Führungskräfte kommunizieren.
  • Wie Teams zusammenarbeiten.
  • Wie sich Organisationen verhalten, wenn Informationen unvollständig sind.

Those factors help shape customer trust just as much as recovery timelines or technical controls do.
And perhaps the most important lesson from Erika and Sam is that trust isn’t earned during easy moments. It’s earned during uncertainty. During ambiguity. During the moments when organizations have to choose transparency over silence and consistency over perfection.

Die ganze Folge ansehen

In this discussion, you’ll discover:

  • Wie Blue Yonder das Vertrauen der Kunden in die Praxis umsetzt.
  • Warum Beständigkeit wichtiger sein kann als sofortige Perfektion.
  • Die Rolle der Kommunikation bei Cybervorfällen.
  • Wie Simulationsübungen die Widerstandsfähigkeit stärken.
  • Warum sich die Rahmenbedingungen für die Cyber-Wiederherstellung in Lieferkettenumgebungen ändern können.

Watch now.

FAQs

Q: Why is trust so important in cyber resilience?

A: Because customers increasingly evaluate organizations based on how they respond during incidents, not simply whether incidents occur.

Q: What does “trust as an operating model” mean?

A: It means trust is continuously reinforced through operational behavior, communication consistency, and transparency – not just during crises.

Q: Why do the first 60 minutes of response matter so much?

A: Early communication can shape customer perception, help reduce uncertainty, and help establish credibility during rapidly evolving situations.

Q: How do tabletop exercises improve resilience?

A: They help teams rehearse coordination, escalation, and communication processes before real incidents occur.

Q: What is the biggest lesson from this discussion?

A: That resilience is typically deeply tied to operational trust, and organizations should build that trust before they need it most.

Q: How can organizations help improve customer trust during incidents?

A: By communicating consistently, prioritizing transparency, and building strong internal coordination long before a crisis begins.

Chris Mierzwa is Senior Director, Portfolio Marketing, at Commvault.

More related posts


Thumbnail_Blog-Architect-for-tomorrow-2026

Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience

Read more about Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience
Thumbnail_Blog-Dangerous-Silos-IDC-Resops-2026 (1)

The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan

Read more about The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan
Thumbnail_Blog-IDC-Resops-2026

Business Continuity Planning for the Cloud-Native Era

Read more about Business Continuity Planning for the Cloud-Native Era

Die wichtigsten Erkenntnisse

  • Cyber-Resilienz in MEDITECH-Umgebungen geht über Datensicherung und -wiederherstellung hinaus; ihr Schwerpunkt liegt auf der Aufrechterhaltung der Patientenversorgung und der Betriebskontinuität bei Störungen.
  • Einrichtungen des Gesundheitswesens sind einem erheblichen Ransomware-Risiko ausgesetzt, weshalb eine schnelle und zuverlässige Wiederherstellung für den klinischen Betrieb unerlässlich ist.
  • Herkömmliche Datenschutzkonzepte tragen den komplexen Abhängigkeiten zwischen klinischen Systemen, Anwendungen und Arbeitsabläufen oft nicht Rechnung.
  • Wirksame Wiederherstellungsstrategien müssen die Wiederherstellung miteinander verbundener Systeme koordinieren, um Ausfallzeiten und betriebliche Beeinträchtigungen zu minimieren.
  • Commvault’s MEDITECH-focused approach helps combine snapshot-based protection, automated recovery workflows, and recovery visibility to strengthen resilience and preparedness.

When ransomware or operational disruption impacts clinical systems, the effects can ripple quickly across the organization, disrupting workflows, slowing staff productivity, and putting timely care delivery at risk. In these moments, the ability to recover quickly and confidently becomes just as important as preventing the disruption in the first place.

That is why Commvault’s approach to protecting MEDITECH environments is centered on recoverability, resilience, and operational readiness, not just data preservation.

Healthcare remains one of the sectors most heavily targeted by ransomware. In a MEDITECH environment, downtime can interrupt medication workflows, delay access to diagnostic information, and force staff into manual workarounds that increase both risk and complexity. In this context, a strategy that looks good on paper is not enough. Health systems need confidence that recovery will perform under real-world pressure.

That is where Commvault can make a meaningful difference.

Warum herkömmlicher Datenschutz für MEDITECH nicht ausreicht

Viele Organisationen stützen sich nach wie vor auf Datenschutzkonzepte, die für allgemeine IT-Umgebungen entwickelt wurden und nicht den betrieblichen Gegebenheiten im Gesundheitswesen Rechnung tragen. Die Wiederherstellung von MEDITECH erfordert ein Verständnis der Abhängigkeiten zwischen den Anwendungen, der Wiederherstellungsreihenfolge, der Validierungs-Checkpoints sowie der Notwendigkeit, klinische Systeme mit minimalen Unterbrechungen wieder in Betrieb zu nehmen.

Eine erfolgreiche Wiederherstellungsstrategie muss mehr umfassen als nur die Wiederherstellung von Daten. Sie muss die koordinierte Wiederherstellung kritischer Systeme, Anwendungen und Arbeitsabläufe unterstützen, auf die sich medizinisches Fachpersonal täglich verlässt. Commvault hilft Unternehmen dabei, diese Komplexität zu bewältigen – mit einer ausfallsicheren Architektur, optimierten Wiederherstellungsabläufen und einem besseren Überblick über die Wiederherstellungsbereitschaft.

Wie Commvault MEDITECH in der Praxis schützt

Commvault’s approach to MEDITECH protection is designed to align with the operational realities of these environments. Rather than relying on a one-size-fits-all backup model, the solution helps organizations capture application-consistent protection points for critical MEDITECH workloads while helping minimize disruption to production operations.

This architecture provides healthcare organizations with a practical path to faster, more confident recovery. Snapshot-based protection can support rapid restoration for operational resilience, while longer-term backup retention strengthens options for audit, compliance, and broader cyber preparedness. The result is a model that supports both day-to-day recoverability and resilience planning for more severe disruption.

Wie sich Cyber-Resilienz in der Praxis gestalten kann

Wenn Ransomware über Nacht zuschlägt

Stellen Sie sich ein regionales Krankenhaus vor, das über Nacht Opfer einer Verschlüsselungsattacke wird. In einem solchen Szenario kann die Möglichkeit, Daten aus unveränderlichen Sicherungskopien wiederherzustellen und einen strukturierten Wiederherstellungsplan umzusetzen, den Unterschied zwischen einem langwierigen Ausfall und einer kontrollierten Wiederherstellung ausmachen. Commvault unterstützt Unternehmen dabei, dieses Risiko zu minimieren – mit sicheren Wiederherstellungsoptionen, die darauf ausgelegt sind, kritische Systeme schnell und fehlerfrei wieder in Betrieb zu nehmen.

Wenn die Backup-Infrastruktur ins Visier gerät

Angreifer versuchen zunehmend, die Backup-Infrastruktur zu kompromittieren, bevor sie Ransomware einsetzen. Daher ist die Widerstandsfähigkeit der Architektur von entscheidender Bedeutung. Mit unveränderlichem Schutz und isolierten Wiederherstellungsoptionen sorgt Commvault dafür, dass saubere Wiederherstellungspunkte auch dann verfügbar bleiben, wenn Angreifer Zugriff auf Produktionssysteme erlangen.

Wann ein Nachweis der Genesung erforderlich ist

Cyberversicherer, Wirtschaftsprüfer und Compliance-Verantwortliche verlangen zunehmend Nachweise dafür, dass die Wiederherstellungsfähigkeiten getestet, dokumentiert und betrieblich einwandfrei sind. Commvault unterstützt diese Bereitschaft durch Validierungsworkflows, Berichterstellung und Nachweise, mit denen Organisationen im Gesundheitswesen ihre Ausfallsicherheit bereits vor dem Eintreten eines Vorfalls unter Beweis stellen können.

Wie Commvault die Ausfallsicherheit von MEDITECH unterstützt

Commvault unterstützt Unternehmen dabei, MEDITECH-Datenbankvolumes mit anwendungskonsistenten Wiederherstellungspunkten zu schützen und schafft so eine solidere Grundlage für die Wiederherstellung, falls klinische Systeme beeinträchtigt werden.

Wiederherstellung unter Berücksichtigung der MEDITECH-Abhängigkeiten

MEDITECH recovery often involves complex relationships between systems and databases. Commvault’s approach helps support coordinated protection of critical workloads and a recovery model designed to help bring systems back online in the appropriate sequence.

Validierte Wiederherstellung mit flexibler Aufbewahrung

Durch die Kombination von Optionen für eine schnelle Wiederherstellung mit einer längerfristigen Aufbewahrung von Backups unterstützt Commvault Teams im Gesundheitswesen dabei, ihre Ausfallsicherheit über das anfängliche Snapshot-Fenster hinaus zu stärken und eine umfassendere Strategie für Wiederherstellungstests, Validierung und Notfallvorsorge zu entwickeln.

Besserer Einblick in die Wiederherstellungsbereitschaft

Eine solide MEDITECH-Resilienzstrategie hängt von operativer Klarheit ab. Commvault unterstützt Teams dabei, Schutz-Workflows zu zentralisieren, den Überblick über die Wiederherstellungsbereitschaft zu verbessern und die Verwaltung kritischer Datensicherungsaufgaben zu vereinfachen.

Einhaltung gesetzlicher Vorschriften und Versicherungsbereitschaft

Von dokumentierten Wiederherstellungsabläufen bis hin zu Aufbewahrungsstrategien, die bei Audits und Compliance-Gesprächen helfen – Commvault unterstützt Organisationen im Gesundheitswesen dabei, ihre Compliance-Dokumentation zu verbessern und eine ausgereiftere Resilienz zu demonstrieren.

Warum die Umsetzung wichtig ist

Eine erfolgreiche Ausfallsicherheit in einer MEDITECH-Umgebung hängt nicht nur von der Auswahl der richtigen platform ab. Sie erfordert auch die Einhaltung validierter Bereitstellungsanforderungen, Infrastrukturkompatibilität und ein Sicherheitskonzept, das der tatsächlichen Funktionsweise von MEDITECH-Systemen in der Praxis Rechnung trägt. Für Organisationen im Gesundheitswesen kann diese Disziplin bei der Umsetzung genauso wichtig sein wie die Wiederherstellungstechnologie selbst. Eine gut durchdachte Resilienzstrategie trägt dazu bei, dass Teams Wiederherstellungsprozesse wie erwartet durchführen können, wenn sie am dringendsten benötigt werden.

Warum gerade jetzt?

Ransomware-Bedrohungen entwickeln sich ständig weiter, und Angreifer nehmen zunehmend die Backup-Infrastruktur ins Visier, bevor sie die Verschlüsselung einsetzen. Gleichzeitig verlangen Cyberversicherer und Compliance-Verantwortliche Nachweise für getestete Wiederherstellungsfähigkeiten und nicht nur für installierte Tools. Für Gesundheitsorganisationen, die MEDITECH einsetzen, ist es dringender denn je, bereits vor dem Eintreten eines Vorfalls Resilienz aufzubauen. Wenn Organisationen heute in die Wiederherstellungsbereitschaft investieren, können sie ihren Betrieb besser schützen, die Wiederherstellung beschleunigen und die Auswirkungen von Störungen in entscheidenden Momenten minimieren.

Im Gesundheitswesen hängt die Bereitschaft zur Wiederherstellung letztlich vom Vertrauen ab: vom Vertrauen darauf, dass kritische Daten geschützt sind, vom Vertrauen darauf, dass Systeme in der richtigen Reihenfolge wiederhergestellt werden können, und vom Vertrauen darauf, dass die Ausfallsicherheit bereits vor dem Eintreten einer Krise getestet wurde. Das ist der Standard, den Commvault Unternehmen in MEDITECH-Umgebungen hilft zu erreichen, und die Grundlage für einen stärkeren, selbstbewussteren Ansatz im Hinblick auf Cyber-Resilienz.

Unternehmen können diese Grundlage weiter stärken, indem sie mit einem auf das Gesundheitswesen spezialisierten Commvault Managed Service Provider zusammenarbeiten. Neben der Technologie selbst erhalten die Teams im Gesundheitswesen Zugang zu Fachwissen, das ihnen dabei helfen kann, Schutzstrategien an die MEDITECH-Anforderungen anzupassen, die Implementierung sicherer zu gestalten und die laufende Betriebsbereitschaft zu verbessern. Für Organisationen im Gesundheitswesen, die sich mit der Komplexität von MEDITECH auseinandersetzen, kann diese Kombination aus ausfallsicherer Technologie und auf das Gesundheitswesen spezialisiertem Fachwissen dazu beitragen, die Vorbereitungen zu beschleunigen und die Ergebnisse bei der Wiederherstellung zu verbessern – gerade dann, wenn es darauf ankommt.

Abschließende Gedanken

Bei der Wiederherstellung in einer MEDITECH-Umgebung geht es um mehr als nur die Wiederinbetriebnahme der Systeme. Es geht darum, die klinischen Arbeitsabläufe wiederherzustellen, auf die sich das Pflegepersonal bei der Patientenversorgung verlässt. Da sich Ransomware-Bedrohungen ständig weiterentwickeln und Organisationen im Gesundheitswesen unter zunehmendem Druck stehen, operative Resilienz unter Beweis zu stellen, darf die Readiness nicht länger als reine Compliance-Maßnahme oder bloße Backup-Strategie betrachtet werden. Organisationen, die Wiederherstellbarkeit, Validierung und Cyber-Resilienz priorisieren, bevor ein Vorfall eintritt, sind besser aufgestellt, um Störungen zu minimieren, die Patientenversorgung zu schützen und im entscheidenden Moment sicher wieder den Betrieb aufzunehmen. Erfahren Sie, wie Commvault Organisationen im Gesundheitswesen dabei unterstützt, die Ausfallsicherheit von MEDITECH zu stärken, die Wiederherstellung zu beschleunigen und mehr Vertrauen in ihre Fähigkeit zu schaffen, Cyberangriffe abzuwehren. Weitere Informationen finden Sie auf unsererMEDITECH-Dokumentationsseite.

FAQs

Q: Why is cyber resilience especially important for MEDITECH environments?

A: MEDITECH environments support critical clinical and operational workflows that directly impact patient care. Cyber resilience helps healthcare organizations recover quickly from disruptions while maintaining essential services and minimizing care delivery interruptions.

Q: How does cyber resilience differ from traditional backup and recovery?

A: Traditional backup and recovery focus primarily on restoring data after an incident. Cyber resilience expands that focus to include operational continuity, rapid recovery, and proactive measures that help reduce the impact of disruptions.

Q: Why are traditional data protection solutions often insufficient for healthcare organizations?

A: Many traditional solutions are designed for general IT environments and may not account for the complex interdependencies between healthcare applications, systems, and workflows. As a result, recovery can be slower and more disruptive.

Q: What challenges do ransomware attacks create for healthcare providers?

A: Ransomware can disrupt access to clinical information, delay care delivery, and increase operational complexity. Healthcare organizations need recovery solutions that enable fast restoration and confidence in recovery outcomes.

Q: How does Commvault support MEDITECH protection and recovery?

A: Commvault’s approach is designed around healthcare operational requirements, providing application-consistent protection, automated recovery processes, and visibility into recovery readiness to help reduce downtime.

Q: What benefits do snapshot-based protection and automated recovery provide?

A: Snapshot-based protection can support faster restoration of critical systems, while automated recovery workflows help streamline recovery efforts. Together, they improve operational resilience and strengthen preparedness for future disruptions.

Chris DiRado is Principal, Product Experience, at Commvault.

More related posts


Thumbnail_Blog-Architect-for-tomorrow-2026

Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience

Read more about Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience
Thumbnail_Blog-Dangerous-Silos-IDC-Resops-2026 (1)

The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan

Read more about The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan
Thumbnail_Blog-IDC-Resops-2026

Business Continuity Planning for the Cloud-Native Era

Read more about Business Continuity Planning for the Cloud-Native Era

When we talk about cyber resilience, the conversation usually centers on technology: tools, platforms, automation. All of that matters. But when something goes wrong, those aren’t the things that determine how well an organization responds.

People are.

In this episode of STRIVE, I sat down with Dr. Jessica Barker, co-CEO of Cygenta and a leading expert on the human and psychological aspects of cybersecurity. I asked her to discuss a part of resilience that doesn’t always get enough attention – the human side.

What happens when pressure rises, when decisions have to be made quickly, and when teams are forced to work together in ways they may not be used to?

Watch the gesamte Folgean, um zu erfahren, was sie zu sagen hatte.

Das Wichtigste auf einen Blick: Was die menschliche Seite offenbart

  • Technology doesn’t fail alone – people and processes are always part of the outcome.
  • Selbstvertrauen unter Druck entsteht durch Vorbereitung, nicht durch Instinkt.
  • Eine klare Zuständigkeit für Entscheidungen trägt dazu bei, das Zögern bei Vorfällen zu verringern.
  • Vertrauen zwischen den Teams trägt dazu bei, die Reaktion und die Wiederherstellung zu beschleunigen.
  • Culture plays a measurable role in resilience – it’s not just tools or architecture.

Wenn der Plan auf die Realität trifft

Every organization has a plan. It’s documented, reviewed, and often approved at the highest levels. But the real test isn’t how that plan reads – it’s how it holds up when people are under pressure.

Because that’s when things change. Decisions don’t always follow the script. Communication isn’t always clean. Priorities shift in real time. And in those moments, resilience becomes less about process and more about behavior.

Vorschau: Cyber-Resilienz als kulturelle Norm

In this moment from the episode, Dr. Barker highlights the importance of aligning cybersecurity with organizational values. Rather than positioning security as a blocker, resilient organizations embed it into culture – as an enabler of productivity, positivity, and business growth.

Die Rolle des Selbstvertrauens

One of the most consistent themes in this conversation is confidence. Not confidence in the tools – confidence in the people using them.

Teams that perform well during incidents aren’t guessing. They’ve seen similar scenarios before. They’ve practiced. They understand how to respond, even when conditions aren’t ideal.

That confidence shows up in small ways with potentially faster decisions, clearer communication, and less second-guessing. And over time, those small differences can add up to a significantly stronger response. 

Entscheidungsfindung unter Druck

When something goes wrong, speed matters – but clarity matters more.

  • Wer darf Entscheidungen treffen?
  • Welche Befugnisse haben sie?
  • Wann sollten sie die Angelegenheit an eine höhere Instanz weiterleiten?

If those answers aren’t clear, teams hesitate. And hesitation creates gaps. One of the most important parts of resilience isn’t just defining processes – it’s defining decision ownership. When people know where they stand, they tend to act faster and with more confidence.

Vertrauen ist der Multiplikator

Technology can help enable better and faster responses, but trust can accelerate them. In most organizations, teams operate in their own lanes. Security focuses on threats, infrastructure focuses on systems, and operations focuses on recovery.

That separation works – until an incident forces everyone together. That’s where trust becomes critical. Teams that trust each other:

  • Informationen freizügiger weitergeben.
  • Arbeiten Sie effektiver zusammen.
  • Konzentrieren Sie sich auf die Ergebnisse statt auf die Zuständigkeit.

Ohne dieses Vertrauen können selbst gut durchdachte Prozesse ins Stocken geraten.

Warum Vorbereitung nach wie vor wichtig ist

It’s easy to assume that strong individuals can carry a response. But even experienced teams rely on preparation, such as tabletop exercises, simulations, cross-team drills, and so on. They’re what build the muscle memory that teams rely on when real incidents occur. Without that preparation, even the most capable teams are forced to improvise.

Die ganze Folge ansehen

In dieser Folge von STRIVE beschäftigen wir uns mit folgenden Themen:

  • Wie menschliches Verhalten die Reaktion auf Vorfälle beeinflusst.
  • Warum klare Entscheidungen unter Druck wichtig sind.
  • Was unterscheidet selbstbewusste Teams von reaktiven?
  • Wie Kultur die Ergebnisse der Genesung beeinflusst.
  • Worauf sich Organisationen konzentrieren sollten, um ihre Widerstandsfähigkeit zu stärken.

Jetzt anschauen.

If you’re thinking about resilience beyond technology, this is a conversation worth your time.

FAQs

Q: Why is the human side of resilience important?

A: Because technology alone doesn’t determine outcomes – people do. Their decisions, communications, and executions under pressure are the keys to success.

Q: What role does preparation play in resilience?

A: Preparation helps build confidence and muscle memory, allowing teams to respond more effectively in real scenarios.

Q: How does trust impact incident response?

A: Trust helps enable faster collaboration, clearer communication, and more efficient decision-making across teams.

Q: Why is decision ownership critical?

A: Without clear ownership, teams hesitate, which can slow response and increase risk.

Q: Can strong tools compensate for weak processes?

A: No. Tools support resilience, but without strong processes and alignment, they can’t deliver effective outcomes.

Q: Where should organizations start improving?

A: Focus on cross-team alignment, clear decision-making structures, and regular scenario-based testing.

Darren Thomsonis Vice President and Chief Technology Officer, EMEA, at Commvault.

More related posts


Thumbnail_Blog-Architect-for-tomorrow-2026

Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience

Read more about Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience
Thumbnail_Blog-Dangerous-Silos-IDC-Resops-2026 (1)

The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan

Read more about The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan
Thumbnail_Blog-IDC-Resops-2026

Business Continuity Planning for the Cloud-Native Era

Read more about Business Continuity Planning for the Cloud-Native Era

For years, recovery planning followed a familiar pattern. Build the plan, document the steps, and assume it will work when needed. For a long time, that approach held up. Hardware failures, isolated outages, even natural disasters – these were scenarios organizations could anticipate and plan for with some level of confidence.
But the equation has changed.
In this episode of STRIVE, I sat down with Commvault’s Jason Cray, Principal Product Experience, to explore a reality we continue to see across organizations of all sizes: Most don’t fail because they lack a recovery plan. They fail because they’ve never proven that plan will hold up under real pressure.
Watch the gesamte Folge.

Die wichtigsten Erkenntnisse: Warum Sanierungspläne scheitern

  • A documented plan isn’t the same as a proven one. If it hasn’t been tested in realistic conditions, it’s still an assumption.
  • Recovery is a team sport. Security, infrastructure, and operations must align – or recovery slows down.
  • Most investment still happens “left of boom.” Prevention matters, but recovery readiness often gets overlooked.
  • Tests decken Schwachstellen auf und stärken das Vertrauen. Ohne sie verlassen sich Unternehmen nur auf Hoffnung.
  • Resilienz ist eine operative Disziplin. Sie erfordert iteratives Vorgehen, Kommunikation und kontinuierliche Verbesserung.

The Problem With ‘It Should Work’

On paper, recovery looks straightforward. You define when to recover to, what needs to come back, and where it should be restored. The process appears logical, structured, and manageable.
But as Jason points out, that simplicity rarely survives real-world conditions.
Plans are written in controlled environments, but they’re executed in chaos. When an incident hits, teams aren’t calmly stepping through documentation – they’re reacting, troubleshooting, and trying to align in real time. That’s where the gap emerges. Not between tools and technology, but between expectation and execution.

Ein kleiner Einblick: Warum Pläne unter Druck scheitern

In this moment from the conversation, Jason and I break down why having a plan isn’t enough – and what it actually takes to know a plan will work when it matters.

We’ve Seen This Before

What’s interesting is that this isn’t a new problem; it’s a familiar one, just in a different context.
If you go back to the early days of disaster recovery, organizations followed a similar pattern. Plans existed, but testing was inconsistent at best. Jason shared an example of spending an entire night helping a client pass a disaster recovery test they thought they were ready for. The plan looked solid. The execution told a different story.
Over time, organizations adapted. They tested more frequently, introduced failover exercises and, in some cases even ran production from secondary environments to prove readiness. That shift from assumption to validation is exactly what Cyber-Resilienz now requires.

Die erste Störung: Kommunikation

If there’s one issue that consistently surfaces, it’s communication.
In many organizations, responsibilities are clearly defined – security handles prevention, infrastructure manages systems, and operations owns recovery. Individually, each team may be doing exactly what they’re supposed to do.
But recovery doesn’t happen in isolation. It depends on how well those teams work together when something goes wrong.
As Jason describes, too often it becomes a handoff model: “We’ve done our part, now it’s someone else’s turn.” That approach introduces delays, confusion, and ultimately risk. During a cyber event, coordination matters more than ownership.

The ‘Left of Boom’ Problem

Another pattern we continue to see is the imbalance in where organizations focus their efforts.
There’s significant investment in prevention – security tools, detection platforms, and defensive strategies designed to stop an attack before it happens. That investment is necessary, and it plays a critical role.
But far less attention is given to what happens after the event.
The assumption is that if enough effort is spent on prevention, recovery becomes a secondary concern. In reality, the opposite is true. At some point, something gets through. And when it does, recovery becomes the defining factor in how an organization responds.

Von der Hoffnung zur Evidenz

This is where the mindset needs to shift.
It’s not about adding more tools or rewriting documentation. It’s about moving from a model based on hope to one grounded in evidence.
Jason highlights a key observation: The organizations that handle disruption well aren’t the ones that avoid incidents – they’re the ones that experience less impact when those incidents occur. They’ve tested their processes. They’ve validated their assumptions. They understand where their gaps are.
Most importantly, they’ve built confidence – not by believing the plan will work, but by proving it.

Fang klein an, baue Schwung auf

For many teams, the challenge isn’t understanding the problem – it’s knowing where to begin.
The answer isn’t to overhaul everything at once. It’s to start small and build from there.
Focus on one or two critical services. Understand what’s required to recover them. Bring together the teams responsible for those systems and test the process end-to-end. From there, expand the scope and continue refining.
This approach does more than improve recovery – it builds alignment, reinforces communication, and creates the foundation for broader resilience.

Die Realität: Kein Plan übersteht den ersten Kontakt

One of the most honest moments in our discussion was this: Even the best plan won’t work exactly as written.
That’s not a failure – it’s expected.
Jason puts it simply: If you don’t have a plan, you will fail. But even if you do have one, it won’t unfold perfectly in the moment.
What matters is how prepared to adapt your teams are. Testing creates that adaptability. It builds the muscle memory needed to respond effectively when conditions don’t match expectations.

Die ganze Folge ansehen

There’s much more we cover in this STRIVE conversation, including:

  • Warum Wiederherstellungspläne oft scheitern, obwohl sie gut dokumentiert sind.
  • Was zeichnet Organisationen aus, die sich effektiv erholen?
  • Wie sich Kommunikationslücken auf die Umsetzung auswirken.
  • Wo soll man bei der Verbesserung der Wiederherstellungsbereitschaft ansetzen?
  • Warum das Testen die Grundlage für Resilienz ist.

Jetzt anschauen.
If you’ve ever questioned whether your recovery plan would actually work, this is a conversation worth your time.

FAQs

Q: Why isn’t having a recovery plan enough?

A: Because most plans are never validated under real-world conditions. Without testing, they remain assumptions rather than proven strategies.

Q: What causes recovery plans to fail?

A: The most common issues in recovery plans are lack of testing, poor cross-team communication, and gaps between documented processes and real execution.

Q: What does “left of boom” mean?

A: Left of boom refers to the focus on preventing incidents before they occur. Many organizations invest heavily here but underinvest in recovery capabilities.

Q: How often should recovery plans be tested?

A: Recovery plans should be tested regularly and under varied conditions. Testing should simulate realistic scenarios, not just controlled exercises.

Q: Where should organizations start?

A: Start with a small set of critical services, align the responsible teams, and test recovery end-to-end before expanding.

Q: What is the key mindset shift?

A: Moving from hope-based planning to evidence-based validation.

Chris Mierzwa is Senior Director, Portfolio Marketing, at Commvault.

More related posts


Thumbnail_Blog-Testing-Once-a-Year-2026

Testing Once a Year Is Not a Resilience Strategy

Read more about Testing Once a Year Is Not a Resilience Strategy
Thumbnail_Blog-IDC-Resops-2026

From Recovery to ResOps™: Building Enterprise Resilience That Scales

Read more about From Recovery to ResOps™: Building Enterprise Resilience That Scales
Readiverse-Featured-Image-888-x-500

Ready Is Good. Resilient Is Better.

Read more about Ready Is Good. Resilient Is Better.

We’ve spent years focusing on identity security in the context of people – who have access, what they can do, and how to control it. That model made sense when most activity in the environment was driven by human users.But that’s no longer the case.Machine identities – applications, services, APIs, and automated workloads – now play a central role in how modern systems operate. They authenticate, communicate, and execute tasks, often without direct oversight. And in many environments, they already outnumber human identities by a wide margin.In this episode of STRIVE, I sit down with Dan Conrad, Principal Technologist and a fellow Field CTO at Commvault. We take a closer look at what that shift means – not just from a security perspective, but also from a governance standpoint. And we explore why so many organizations are still treating this as a secondary concern.Watch the gesamte Folge.

Das Wichtigste auf einen Blick: Wohin sich das Risiko verlagert

  • Die Zahl der Maschinenidentitäten wächst schneller als die der menschlichen Identitäten, oft um ganze Größenordnungen.
  • Governance models haven’t kept pace, creating blind spots in access and control.
  • Visibility is the core challenge. Many teams don’t fully understand how machine identities behave.
  • Die Ausweitung von Zugriffsrechten geht über die Benutzer hinaus, da Maschinenidentitäten häufig über dauerhaften Zugriff verfügen.
  • Resilienz hängt davon ab, dass man den Aufgabenbereich dieser Maschinenidentitäten versteht und steuert, bevor sie zu einem Problem werden.

Das Identitätsmodell hat sich verändert

For a long time, identity management was relatively straightforward. You could map users to roles, define access policies, and build controls around predictable behavior. Even with complexity, the model was still anchored in human activity.Machine identities have broken that model.They’re created dynamically, often as part of development or deployment processes. They interact across systems in ways that aren’t always visible or well-documented or audited. And unlike human users, they don’t follow a clean lifecycle – they aren’t onboarded and offboarded in the same structured way.That creates a different kind of challenge. It’s not just about controlling access anymore. It’s about understanding how that access is being used, how it evolves, and how it connects across the environment.

Sneak Peek: You Can’t Phish a Non-Human Identity

In this moment from the STRIVE discussion, Dan describes how attackers aren’t targeting nicht-menschliche Identitäten directly through phishing – they’re malicious actors using compromised human accounts through social engineering, as a steppingstone to escalate privileges and impersonate powerful machine identities. Once inside, techniques like pass-the-hash and overprivileged service accounts allow attackers to move laterally and vertically, even after passwords are reset.

Die Governance-Lücke

The real issue isn’t that machine identities exist , it’s how they’re governed.In most organizations, there’s a clear process for managing human access:

  • Die Anträge werden genehmigt.
  • Die Berechtigungen werden überprüft.
  • Änderungen werden nachverfolgt.

There’s a level of discipline that comes from years of focus on user identity. However, machine identities often fall outside of that structure. They’re created quickly to support applications or automation. They’re granted the permissions needed to function, sometimes more than necessary. And over time, those permissions persist. These overprovisioned accesses are rarely audited, reviewed, and more importantly rarely reduced.That’s where the gap forms.It becomes difficult to answer basic questions about access. Not because the information doesn’t exist, but because it hasn’t been organized or managed in a way that makes it usable.

Transparenz vor Kontrolle

When organizations start to address this problem, the instinct is often to tighten controls.

  • Berechtigungen einschränken
  • Zugriff einschränken
  • Neue Richtlinien anwenden

But control without visibility doesn’t solve much.If you don’t understand how identities are being used, the business context of it in terms of  where they connect, what they interact with, and how they move across systems, then any attempt to restrict them becomes reactive and could result in slowing down business operations.That’s why visibility needs to come first.Once you can see how machine identities behave, patterns start to emerge. You can begin to understand where access is excessive, where dependencies exist, and where risk is concentrated. From there, governance can become more precise and more effective.

Ein etwas anderes Privilegienproblem

Privilege sprawl isn’t new. Most organizations have spent years trying to manage excessive access among human users.Machine identities introduce a similar issue, but with a different dynamic. Their access is often embedded into systems. It’s persistent, automated, and rarely questioned once it’s in place. That makes it harder to detect and easier to overlook.And when something goes wrong, those identities can become a pathway for malicious actors to exploit

Wo soll man anfangen?

For most organizations, the challenge isn’t awareness, it’s knowing where to start.The first step isn’t a major transformation. It’s building clarity. Understanding how many machine identities exist. Where they’re being created. What permissions they have. How they’re used. And most importantly, confirming that a human user is mapped to a collection of nicht-menschliche Identitäten for the purposes of auditability and accountability.Those questions sound simple, but they’re often difficult to answer. And that’s exactly why they matter. Because once you can answer them, you’re no longer operating in the dark.

Die ganze Folge ansehen

In this installment of STRIVE, we go deeper into how machine identities are changing the way organizations should think about access, governance, and resilience. It’s a practical conversation about what’s happening now – and what needs to change moving forward.Jetzt anschauen.

Ressource

If you’re interested in learning more about this topic, check out this e-book on nicht-menschliche Identitäten.

FAQs

Q: What is a machine identity?

A: A machine identity is a non-human identity used by applications, services, or systems to authenticate and interact with other resources.

Q: Why are machine identities becoming a bigger risk?

A: Because they are increasing in number, often have persistent access, and are not always governed as strictly as human users.

Q: How are they different from user identities?

A: They operate continuously, are embedded in automated workflows, and often lack structured lifecycle management.

Q: What is the biggest challenge organizations face in governing nicht-menschliche Identitäten?

A: Visibility. Many teams don’t have a clear understanding of how many machine identities are created, used, or interconnected.

Q: How does this impact resilience?

A: If compromised, machine identities can enable a malicious actor’s rapid movement across systems, making incidents harder to contain and recover from.

Q: Where should organizations start?

A: By identifying machine identities, understanding their permissions, and building governance practices that match their scale and complexity. And most importantly, confirming that a human user is mapped to a collection of nicht-menschliche Identitäten for the purposes of auditability and accountability

Vidya Shankaran is Field CTO at Commvault.

More related posts


Thumbnail_Blog_Identity-Resilience-Vishing_2026

Are You Ready for the Industrialized Vishing Attack?

Read more about Are You Ready for the Industrialized Vishing Attack?
Thumbnail_Blog-Identity-Resilience-MachineID-2026-Linkedin

The Machine Identity Blind Spot Is Now a Primary Attack Surface

Read more about The Machine Identity Blind Spot Is Now a Primary Attack Surface
Thumbnail_Blog-Help-Desk-2026-Linkedin

When the Help Desk Becomes the Front Door to Your Entire Network

Read more about When the Help Desk Becomes the Front Door to Your Entire Network

Jahrzehntelang lag der Schwerpunkt im IT-Betrieb auf der Verfügbarkeit:

  • Sorgen Sie dafür, dass die Infrastruktur weiterläuft.
  • Erreichen Sie Ihr Wiederherstellungszeitziel (RTO).
  • Erfüllen Sie Ihr Wiederherstellungsziel (RPO).

But modern cyber threats don’t respect infrastructure boundaries – and recovery isn’t just about restoring systems anymore. It’s about restoring clean, trusted data – across teams, under pressure.
In this episode of STRIVE, I sat down with Stephen Foskett, founder and president of the Futurum Group’s Tech Field Day, to discuss an emerging discipline: resilience operations – or ResOps.

And it’s more than a buzzword. It’s a shift in how organizations think about recovery intelligence.
Watch the gesamte Folge.

Das Wichtigste auf einen Blick: Was sich bei ResOps ändert

  • ResOps moves recovery from infrastructure-focused to business-focused. It’s not just about bringing systems back online – it’s about restoring trusted, usable data.
  • Traditional RTO and RPO metrics aren’t enough anymore. Die „Mean Time to Clean Recovery“(MTCR) entwickelt sich zu einer aussagekräftigeren Methode zur Messung der Ausfallsicherheit.
  • Breaking down silos is foundational to cyber readiness. Security, infrastructure, and DevOps must operate in sync – not in parallel.
  • Resilience is an operational discipline, not a tool. Culture, communication, and coordination matter as much as technology.
  • Recovery intelligence is becoming a competitive differentiator. Organizations that recover cleanly and quickly protect revenue, reputation, and trust. 

From IT Ops to ResOps: What’s Changed?

Stephen blickt auf eine frühere Ära der IT zurück, in der Teams häufig Systeme betreuten, ohne die damit verbundenen Geschäftsanwendungen vollständig zu verstehen. Unter „Wiederherstellung“ verstand man damals die Wiederherstellung der Infrastruktur. Heute reicht dieses Modell nicht mehr aus. Moderne Umgebungen zeichnen sich durch folgende Merkmale aus:

  • Verteilt
  • Cloud
  • DevOps-orientiert
  • Sicherheitsrelevant
  • Eng mit den Einnahmequellen verzahnt

ResOps acknowledges that recovery is no longer an isolated IT function. It’s a cross-functional discipline that helps connect infrastructure, software development, and security with real business outcomes.

Why Traditional Metrics Don’t Tell the Whole Story

RTO. RPO. These metrics have guided disaster recovery planning for years. But as Stephen explains, restoring quickly isn’t enough if the data you restore isn’t clean.
Enter a more meaningful metric: MTCR. It’s not just how fast you recover; it’s how fast you can recover to a verified, trusted state.
In a ransomware event, that difference matters enormously. Restoring compromised data can restart an attack cycle. ResOps focuses on restoring operational integrity – not just functionality.

Vorschau: Warum eine „Clean Recovery“ wichtig ist

In this moment from STRIVE, Stephen explains why traditional recovery metrics miss the mark – and why recovery is a cross-functional discipline.

Das eigentliche Hindernis: Organisatorische Silos

Technology isn’t usually the biggest blocker to resilience. Structure is. Security teams often report to one executive. Infrastructure teams to another. Application teams to yet another. Each with different priorities, different incentives, and different definitions of success.
ResOps challenges that fragmentation.
Stephen discusses how collaborative workshops and cross-functional alignment are helping break down those silos. Because during a cyber event, organizational misalignment slows recovery more than tooling gaps ever will.

Warum Commvault sich in dieser Diskussion engagiert

STRIVE isn’t about product features. It’s about how recovery thinking is evolving. ResOpsdeckt sich genau mit dem, was wir in der Praxis beobachten:

  • Kunden, die bei Vorfällen Schwierigkeiten mit der Koordination haben.
  • Unternehmen, die ihre Infrastruktur wiederherstellen, aber die Datenintegrität in Frage stellen.
  • Die Unternehmensleitung fordert Kennzahlen, die die tatsächlichen geschäftlichen Auswirkungen widerspiegeln.

The concept of MTCR reframes recovery intelligence around business trust – and that’s where the industry is heading. Recovery is no longer a back-office process. It’s an executive concern.

Die Zukunft der Recovery Intelligence

Looking ahead, ResOps is likely to mature rapidly. Over the next 12–18 months, organizations are expected to:

  • Sicherheits- und Wiederherstellungsabläufe enger miteinander verknüpfen.
  • Führen Sie neue, auf die Erholung ausgerichtete Kennzahlen ein.
  • Resilienz bereits in früheren Phasen des Anwendungslebenszyklus umsetzen.
  • Investieren Sie in eine Lösung, die saubere Daten von kompromittierten Daten unterscheidet.

Cyber threats are accelerating. Recovery strategies must evolve at the same pace. ResOps helps provide a framework for doing that.

Die ganze Folge ansehen

In dieser Folge sprechen wir über:

  • Wie sich ResOps vom herkömmlichen IT-Betrieb unterscheidet.
  • Warum MTCR dazu beiträgt, die Kennzahlen zur wirtschaftlichen Erholung neu zu definieren.
  • So sieht organisatorische Ausrichtung in der Praxis aus.
  • Wie die DevOps-Kultur die Ausfallsicherheit beeinflusst.
  • Wohin sich die Recovery-Intelligence voraussichtlich als Nächstes entwickeln wird.

Jetzt anschauen.
If you’re responsible for cyber readiness, continuity, or recovery strategy, this is a must-watch discussion.

FAQs

Q: What is ResOps?

A: ResOps (Resilience Operations) is an emerging discipline that integrates IT operations, security, DevOps, and business stakeholders to help improve recovery intelligence and organizational resilience.

Q: How is ResOps different from traditional IT operations?

A: Traditional IT ops focuses primarily on infrastructure uptime. ResOps expands that focus to include clean data recovery, cross-functional coordination, and business alignment.

Q: What is Die „Mean Time to Clean Recovery“ (MTCR)?

A: MTCR measures how quickly an organization can restore verified, clean data and resume safe operations after a cyber event – not just how quickly systems are brought back online.

Q: Why are metrics like RTO and RPO insufficient in modern environments?

A: They measure speed and data currency, but not data integrity. In ransomware scenarios, restoring compromised data can extend disruption.

Q: How can organizations start implementing ResOps?

A: Begin by:

    • Abstimmung zwischen Sicherheits-, Infrastruktur- und DevOps-Teams.
    • Bewertung von Wiederherstellungskennzahlen über RTO/RPO hinaus.
    • Testen von sauberen Wiederherstellungsprozessen.
    • Abbau operativer Silos.
    • Das Resilienzdenken bereits in einer früheren Phase der Systemgestaltung einbeziehen.

Q: Why is recovery intelligence becoming more important?

A: As cyber threats grow more sophisticated, the ability to recover cleanly, quickly, and confidently directly impacts revenue, customer trust, and regulatory posture.

Darren Thomsonis a Field CTO at Commvault.

More related posts


Thumbnail_Blog-Testing-Once-a-Year-2026

Testing Once a Year Is Not a Resilience Strategy

Read more about Testing Once a Year Is Not a Resilience Strategy
Thumbnail_Blog-IDC-Resops-2026

From Recovery to ResOps™: Building Enterprise Resilience That Scales

Read more about From Recovery to ResOps™: Building Enterprise Resilience That Scales
Readiverse-Featured-Image-888-x-500

Ready Is Good. Resilient Is Better.

Read more about Ready Is Good. Resilient Is Better.

Die wichtigsten Erkenntnisse

  • Frontier AI is collapsing vulnerability remediation windows,prevention alone can no longer guarantee security.
  • The question that boards,regulators,and insurers are now asking is not “Do we have backups?” but “Can we prove we can recover cleanly?”
  • Backups are not recovery: a copy tells you data exists,not whether it is clean or restorable.
  • Mean Time to Clean Recovery (MTCR) must become a board-level,continuously measured number – not a theoretical estimate.
  • An Isolated Recovery Environment – air-gapped,immutable,hardened,and identity-isolated – is the baseline,not an advanced capability.
  • What counts as “clean” will keep changing as AI models grow more capable of finding compromises humans cannot anticipate.

I have spent a large part of my career running production systems. I know backup environments from the inside,the ones customers actually trust. I know recovery plans as the things that only reveal their weaknesses when something has already gone wrong. That experience changes how you think about cyber resilience.
From a distance,backup and recovery sounds manageable. Protect the data,store copies,document the runbook,test when you can,restore when you need to. But anyone who has run these environments at scale knows the harder truth: Recovery is where assumptions go to be tested. And right now,too many organizations are operating on assumptions that no longer fit.
For years,security operated on a familiar sequence: Find the vulnerability,patch it,harden the environment,monitor for activity. That model still matters. But the window it depends on is collapsing.
Frontier AI has changed the velocity of vulnerability discovery,attack path chaining,and exploit generation. Models like Claude Mythos and GPT-5.5-Cyber have already demonstrated what this looks like,so far in controlled,early-access testing that still relied on human expertise and carried meaningful false-positive rates,but the trajectory is unmistakable. As access widens,the same capability moves into attackers’ hands.
In a single month,Palo Alto Networksdisclosed 26 CVEs,representing 75 underlying issues,after adopting frontier AI models for code scanning,compared with its typical volume of fewer than five CVEs per month.
Researchers also are warning that AI-assisted discovery is collapsing remediation windows,with someexploits now emerging within minutes of disclosure. When the patch window disappears,the remediation math stops working. Prevention cannot carry the full weight of readiness.
Prevention still matters,but it no longer defines readiness. The customers I talk to are not asking whether they need more controls. They already know they do. They are asking whether their business can recover cleanly when those controls fail,when attackers move faster than remediation cycles,or when compromise has been present longer than anyone realized.
That question is now what boards,regulators,and insurers are forcing. They have moved past “Do we have backups?” and toward something more consequential: “Can we prove we can recover cleanly?”

That proof starts with one distinction most organizations still get wrong: Backups are not recovery.
A backup tells you a copy exists. It does not tell you whether the data is clean,whether application dependencies are intact,whether identity services can be safely restored,or whether the recovery sequence still reflects the current environment.
I have reviewed plans that looked complete until someone tried to execute them. The runbook was there,but outdated. The restore worked but took three times longer than the estimate. The system came back,but downstream applications could not connect. None of that is unusual. It is exactly what real testing is supposed to surface. The problem is most organizations discover these gaps during an actual incident.
The metric that matters most when something goes wrong is how quickly you can return to a known-good state. That is why Mean Time to Clean Recovery (MTCR)needs to become a board-level number,not a theoretical estimate in a plan,but a measured,validated time.

The Moving Target: What’s Clean Today May Not Be Clean Tomorrow

With Frontier AI models,the honest answer is this: you cannot guarantee that every vulnerability will be found and remediated in time. Attackers leveraging the same models are discovering and chaining exploits faster than any remediation program can realistically keep pace with. That is not a failure of your security team. It is the new physics of the threat landscape.
What you can control is your ability to recover. That means an Isolated Recovery Environment – backups air-gapped from the internet,unreachable from the production network,and protected from the lateral movement that defines a sophisticated breach. It means immutability and compliance lock,so no credential,however privileged,can shorten retention or delete data outside an authorized process. And it means ResOps in practice: not just backing up data,but continuously testing recovery,automating integrity validation,and measuring your MTCR – the validated time to return to a known-good state.
But here is the part most organizations are not yet accounting for: what counts as “clean” is not a fixed line. As AI models grow more capable,they will increasingly find vulnerabilities that the human mind simply cannot anticipate,novel attack paths,dormant implants,subtle corruptions embedded long before detection. A recovery point that is clean by today’s standards may carry compromise that tomorrow’s AI-assisted forensics will surface. That means your definition of clean must evolve continuously. MTCR is not a number you set once. It is a discipline you maintain,revisiting what clean means,updating your validation criteria,and treating resilience as a living standard rather than a certification you pass once.
So what is a good MTCR? Based on what I have seen work in practice,the target for your entire minimum viable company – the smallest set of systems that lets you keep operating,which I define precisely below,should be under six hours. Six hours is achievable with the right architecture: an IRE ready to run,a pre-validated recovery sequence,and runbooks that are executable rather than readable. If your current MTCR is measured in days,the gap is almost always one of those three.

Four Steps to Stay Resilient in the Frontier AI Era

Accepting that prevention alone is not enough is the starting point. From there,the work gets specific. Here is where I tell organizations to focus.

1. Evaluate your actual recovery risks.

Most recovery risk assessments ask the wrong questions. “Do backups exist?” is not the same as “Can we recover cleanly?” The harder questions are: Can critical systems be restored without reintroducing the threat? Are recovery environments isolated from compromised production systems? Are recovery plans mapped to current dependencies – not the architecture from two years ago?

In a fast-moving vulnerability environment,the gap between “we have backups” and “we can recover” is where organizations get hurt. Assessing that gap honestly,before an incident forces the issue,is where resilience planning must start. That assessment needs to include a business impact analysis: which systems have a recovery window measured in minutes,which in hours,and which can wait a day. Without that tiering,every system looks equally urgent during an incident,and nothing gets restored fast enough.

2. Make isolated recovery and air gapping the baseline – not the exception.

If you are still treatingair-gapped,immutable copies as an advanced capability rather than a standard requirement,that assumption no longer holds. When exploitation timelines compress to minutes,you need fallback options that are structurally separated from production identity,network,and management planes – logically or physically isolated,immutable,and with no live path back to production that an attacker can follow.
The goal is not just protection from the current threat,but maintaining clean recovery options when a vulnerability you have not patched yet gets exploited. That happens now. Plan for it.
Isolation only holds if the infrastructure around it is hardened. That means backup infrastructure on hardened operating systems,not generic images,and ideally on physical servers that survive a hypervisor-layer attack. It means encryption keys stored outside the backup platform,in an external vault with just-in-time access and no dependency on production Active Directory. And it means treating your backup domain as a separate identity boundary: no trust to production AD,mandatory MFA,and multi-person authorization for destructive operations. None of this is exotic,it is the baseline for your environment to recover into an uncompromised space.
Equally important is the question of what you are recovering from. Industry incident-response data consistently puts median breach dwell time in the range of weeks,not days. That means your recovery copies need to reach back far enough to find a genuinely clean point,not just yesterday’s backup. Critical systems warrant multiple geographically separated copies,including at least one immutable copy and one that is fully offline. Retention policy is not a storage cost decision. It is a security decision.

3. Know which systems the business cannot operate without – and recover those first.

Most organizations discover their recovery sequence during an incident. That’s why the first 24–48 hours aren’t spent restoring systems,they’re spent deciding what matters.
Organizations know they have to recover identity platforms,billing systems,operational databases,and core infrastructure. What they often have not mapped is the order,the dependencies between those systems,and the downstream applications that cannot function until specific services are back.
This gets more complex as AI becomes embedded in business operations. Data pipelines,model repositories,vector databases,agentic workflows – these are now operational dependencies,not just technical infrastructure. If your recovery sequencing does not account for them,your recovery time estimates are probably wrong.
Defining what it means to operate as a minimum viable company (the smallest set of systems required to keep the business running) and building recovery around that definition is not a theoretical exercise. It is the practical answer to the question every executive team will ask during an incident: What do we bring back first?

In my experience helping customers through active incidents,the first 12 hours answer that question whether you have planned for it or not – what gets recovered in that window becomes your MVC by default. The organizations that come through fastest decided in advance: they knew exactly which systems had to be back within 12 hours and had validated they could do it. If your MVC does not fit in 12 hours,it is not your MVC,it is a wish list. The work is to keep trimming until what remains can realistically be restored in that window,then test it until you can prove it.

4. Automate resilience and test continuously – not on a calendar schedule.

A recovery plan that lives in a document and is reviewed annually is not a recovery capability. It is a hypothesis that has never been tested against reality.
The problem with calendar-based testing is what it misses between cycles. Environments change constantly: new workloads,updated dependencies,infrastructure that has drifted from what the runbook describes. By the time the annual test runs,it is validating a snapshot of an environment that no longer exists. In a threat landscape where exploitation can happen within minutes of disclosure,that lag is not acceptable.Threat scanning,clean recovery point identification,dependency-aware restoration,and recovery orchestration all need to be automated and running continuously. Not because automation is a best practice,but because the manual alternative cannot keep pace with how fast things now move.
Continuous testing also depends on continuous detection. You cannot select a clean recovery point if you do not know when the compromise began. That is why threat detection,anomaly scanning of backup data,and recovery-point analysis have to feed each other: detection tells you which copies predate the intrusion,and that determination drives which point you actually recover from. Without that link,you are restoring to a date you hope is clean rather than one you have verified,and in a Frontier AI threat landscape,hope is not a recovery strategy.
What continuous testing surfaces is different from what annual tests find. Calendar tests tend to confirm the plan works under controlled conditions. Continuous testing finds the dependency that changed last month,the recovery sequence that breaks when a specific workload is added,the identity service that takes twice as long to restore as the estimate assumed.
Those are the gaps that matter during a real event,and the only way to find them before an incident does is to be testing all the time.
Testing also needs to happen in the right environment. A recovery test that runs against production infrastructure does not tell you whether you can recover when production is compromised. Cleanroom testing – validating restoration in a fully isolated environment with no connectivity back to production – is how you confirm your backup copies are genuinely usable under incident conditions. That includes recovering identity services,external key management,and Tier 0 applications in isolation,with dedicated break-glass accounts that exist outside your normal directory.
What makes daily testing viable is validate restore,a recovery type that exercises the full restore path for every critical asset without touching production. Your backup platform needs to support this natively; if it cannot run an automated,non-disruptive recoverability test across your MVC every day,you do not actually know whether your backups work. In Commvault,this restores against your critical-asset groups,with automated reporting on the recovery status of every protected system.
The same applies to your runbooks. A runbook that lives in a Word document or PDF is a reference manual,not an operational tool – it assumes someone has the time,clarity,and access to read it under pressure. Real runbooks are digital scripts that execute the recovery sequence and validate each step,confirming the application actually works before moving on: not the service started” but “the application responded correctly to a synthetic transaction.” Commvault’s Cleanroom Runbooks are built for this – executable workflows that drive an end-to-end recovery in an isolated environment without a human interpreting a document at every step.
One final point that rarely makes it into recovery plans until it is too late: during a serious incident,your corporate communications infrastructure may itself be compromised or unavailable. Email,Teams,and Slack run on the same infrastructure attackers target. Know in advance which out-of-band channels your team will use to coordinate,and make sure those channels are tested alongside your technical recovery procedures.
Hear more from Commvault’s Chief Security Officer Bill O’Connell on the four critical steps for resiliency in the AI era.

Resilience Is an Operating Discipline,Not a Project

The organizations that will hold up under frontier AI-accelerated threats are the ones that treat resilience as anoperating discipline — measured MTCR,continuous validation,and a recovery capability they have proven,not assumed.
The problem isn’t that attacks are getting faster. It’s that recovery hasn’t caught up,and until it does,the math doesn’t work.


FAQs

Q: What is Mean Time to Clean Recovery (MTCR) and why does it matter?
A: MTCR measures how quickly an organization can return to a verified,known-good state after a cyberattack – not just restore data,but confirm it is clean and that application dependencies are intact. It should be a board-level metric with a measured,validated time,not a theoretical estimate buried in a recovery plan. The target for a well-architected MVC – covering all identity systems,critical applications,and isolated environment readiness – is under six hours.
Q: What is an Isolated Recovery Environment and how is it different from a standard backup?
A: An Isolated Recovery Environment is a fully air-gapped,immutable copy of critical data that is structurally separated from production networks,identity systems,and management planes. A standard backup tells you a copy exists. An IRE tells you that copy is protected from the same attack that hit your production environment.
Q: How do we know whether we can actually recover today?
A: The only honest answer comes from testing,not documentation. If you cannot point to a recent,validated recovery of your minimum viable company – ideally a daily automated test – then you do not know,you are assuming. A defensible answer to the board is a measured MTCR backed by continuous validation,not a recovery plan that looks complete on paper.
Q: What do regulators and cyber insurers now expect?
A: The bar has moved from “Do you have backups?” to “Can you prove you can recover cleanly,and how fast?” Regulators increasingly expect demonstrable recovery capability and tested resilience; insurers increasingly price coverage – and pay claims – based on evidence of isolated,immutable backups and validated recovery times. A measured MTCR and a documented testing cadence are becoming table stakes for both.
Rajiv Kottomtharayil is Chief Product Officer at Commvault.

More related posts


Thumbnail_Blog-Architect-for-tomorrow-2026

Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience

Read more about Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience
Thumbnail_Blog-Dangerous-Silos-IDC-Resops-2026 (1)

The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan

Read more about The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan
Thumbnail_Blog-IDC-Resops-2026

Business Continuity Planning for the Cloud-Native Era

Read more about Business Continuity Planning for the Cloud-Native Era

Blog

Schutz von KI-Workloads: Wie können Unternehmen im Zeitalter der KI Ausfallsicherheit gewährleisten?

AI resilience helps enable protection, recovery, and governance of AI workloads, data, and models by combining threat detection, clean recovery, and controlled data access.

Häufig gestellte Fragen

Was versteht man unter KI-Resilienz?

AI resilience is the ability to protect, recover, and govern AI systems across their full lifecycle. Commvault’s Protect and Leverage AI capabilities help verify that data, models, and pipelines remain secure, recoverable, and trustworthy — even when disrupted by cyber threats, failures, or operational complexity in hybrid and multi-cloud environments.

Warum ist der Schutz von KI-Workloads wichtig?

KI-Workloads basieren auf verteilten Daten, Modellen und Infrastrukturen, wodurch sie anfällig für Bedrohungen wie Datenvergiftung und Modellbeschädigung sind. Ihr Schutz trägt dazu bei, die Datenintegrität zu wahren, Betriebsrisiken zu reduzieren und das Vertrauen in KI-gestützte Geschäftsprozesse aufrechtzuerhalten. Commvault hilft bei der Bewältigung dieser Herausforderungen mit Metallic AI, das ML-gesteuerte Erkennung, geführte Recovery und Automatisierung in der gesamten Commvault Cloud vereint.

Was umfasst der umfassende Schutz des gesamten KI-Stacks?

Full AI stack protection safeguards data pipelines, vector databases, models, metadata, configurations, and compute infrastructure. Commvault Cloud Unity covers this breadth — including unified data platforms like Amazon Redshift and Google BigQuery, vector retrieval systems, and compute infrastructure — enabling complete and consistent recovery of AI workloads across hybrid and multi-cloud environments.

Warum ist eine saubere Wiederherstellung in KI-Umgebungen so wichtig?

Clean recovery confirms that restored data is free from corruption, malware, or inconsistencies. In AI systems, compromised data leads to inaccurate outputs and biased decisions. Commvault Synthetic Recovery addresses this by analyzing multiple backup versions to assemble a validated recovery point — so restored AI workloads produce trusted, accurate outputs.

Inwiefern trägt KI zur Verbesserung des Datenschutzes und der Betriebsabläufe bei?

Commvault embeds AI across the protection lifecycle — automating threat detection, optimizing backup scheduling, and predicting storage needs through ML-enabled capabilities. Arlie, Commvault’s AI assistant, improves user experience through natural language interactions, guided workflows, and intelligent insights, helping security and IT teams manage complex AI environments more efficiently.

Was versteht man unter verantwortungsvoller KI im Datenschutz?

Verantwortungsbewusste KI ermöglicht es Systemen, transparent, geregelt und kontrolliert zu arbeiten. Commvault unterstützt dies durchDaten aktivieren — a governed workspace that applies encryption, immutability, and role-based access controls to curate and extend trusted data to AI and analytics platforms, helping prevent misuse and maintain compliance while enabling innovation.


Die wichtigsten Erkenntnisse

  • Mythos wird die Erkennung von Sicherheitslücken wahrscheinlich in einem Ausmaß und mit einer Geschwindigkeit beschleunigen, die herkömmliche, von Menschen gesteuerte Behebungsabläufe übertreffen.
  • Grundlegende Sicherheitsmaßnahmen wie das Installieren von Patches, Backups in isolierten Netzwerken und ein konsequentes Schwachstellenmanagement sind nach wie vor von entscheidender Bedeutung, reichen jedoch allein möglicherweise nicht mehr aus.
  • Die größte Herausforderung besteht darin, den Schwerpunkt von der Erkennung auf die Handlungsfähigkeit zu verlagern, da das Ausmaß der Sicherheitslücken die derzeitigen operativen Grenzen sprengt.
  • AI resilience depends on the ability to recover coherent systems – not just data – across models, pipelines, and permissions.
  • Unternehmen, die sich in dieser frühen Phase proaktiv anpassen, dürften deutlich besser aufgestellt sein als solche, die mit Maßnahmen zögern.

A few weeks ago, I was in a room with a group of CIOs and CISOs when the conversation turned to Mythos and Project Glasswing. The energy was immediate – these are people who have lived through a lot of hype cycles, and this commanded their attention.

The reactions landed in two camps. One: The threat categories aren’t new – organizations with solid vulnerability management and trusted air-gapped backups will be better positioned than those without. Two: The velocity is different – not just what Mythos can find, but how fast, how fast bad actors could leverage AI for machine-speed attacks, and what that does to the math that most vulnerability management programs are built on.

Both were right. That’s what made the conversation worth writing about.

What Mythos Changes – And What It Doesn’t

Mythos is Anthropic’s AI model for autonomous vulnerability discovery. It can find and chain critical exploits across major operating systems at a success rate that is believed to have no real precedent in this domain.

Project Glasswing – the consortium of companies brought in to test and harden their systems before Mythos or similar capabilities reach adversaries – is the signal that this is real, it is here, and the window for getting ahead of it is short.

The fundamentals-first view holds: patching matters, virtually air-gapped backups matter, vulnerability management discipline matters. None of that changes with Mythos. What changes is the production rate on the other side of those programs.

The question after Glasswing isn’t whether you have a vulnerability management program. It’s whether it was built for findings that arrive in a trickle – or a tsunami.

Most programs were built for the trickle. Periodic assessments, CVSS-based prioritization queues, patch and testing cycles measured in weeks. That cadence made sense when the pace of discovery matched the pace of human-led processes. Mythos-class capability breaks that assumption – the volume of exploitable findings may exceed what most organizations can process through the workflows they have today.

The issue isn’t detection. It’s capacity to act – and what happens when the gap between discovery and remediation widens faster than you can close it.

tröpfchenweisen Zustrom

When prevention timelines are compressed, the resilience question moves to the front of the line. If you can’t guarantee you’ll patch everything before something is exploited – and increasingly, you can’t – the questions that matter shift: How fast do you detect? How do you contain? And when you recover, what exactly are you recovering to?

That last question is harder than it sounds, especially for organizations with progressive agentic interactions. An AI system isn’t just data. It’s a model version, a training pipeline, a vector database, a set of agent identities and permissions – all of which need to reflect the same operational state to constitute something you can actually trust.

Most organizations can restore individual components. Very few can prove that what they’ve restored is coherent.

Recovering an AI system isn’t a data restoration problem. It’s a coherence problem – and the gap between those two things is where most enterprises are currently exposed.

This is the thread that connects Mythos to the broader AI resilience conversation. It isn’t that Mythos introduces a new type of risk that requires a new framework.

It’s that Mythos compresses the timeline in a way that surfaces existing gaps faster, with less runway to close them before something goes wrong, thereby increasing the change that something will go wrong before an organization can properly remediate vulnerabilities.

The Window Is Open. It Won’t Stay That Way.

Glasswing was designed to give defenders a head start. The organizations that use this window deliberately – stress-testing their vulnerability programs for volume, getting AI resilience infrastructure to a state they can defend, and treating recovery as something that has to be provable before an incident, not assembled during one – will be in a materially better position than those that wait.

The fundamentals still apply. The urgency is new.

„The Agentic Enterprise: Warum KI-Resilienz ein System of Record erfordert“ – Commvault’s latest Readiness Report – examines the AI resilience infrastructure gaps that determine whether organizations can answer the hard recovery questions when the pace of threats demands it.

FAQs

Q: What is Mythos and why is it significant?

A: Mythos is an AI model designed for autonomous vulnerability discovery, capable of identifying and chaining exploits across systems at unprecedented speed. Its significance lies in how it compresses the timeline between vulnerability discovery and potential exploitation, raising the stakes for defenders.

Q: Does Mythos change the fundamentals of cybersecurity?

A: No, core practices like patching, backups, and vulnerability management still matter. What is changing is the volume and velocity of threats, which puts pressure on existing processes that were designed for slower, more predictable workflows.

Q: Why may current vulnerability management programs struggle?

A: Many programs were built for a steady flow of findings, not the surge enabled by AI-driven discovery. As a result, organizations face a growing gap between identifying vulnerabilities and actually remediating them.

Q: What does “resilience” mean in the context of AI systems?

A: Resilience goes beyond restoring data – it involves recovering an entire AI system in a coherent, trustworthy state. This includes models, training pipelines, vector databases, and access controls all aligning correctly.

Q: Why is recovery becoming more important than prevention?

A: As prevention timelines shrink due to faster exploitation, it is becoming unrealistic to patch everything in time. This shifts focus to how quickly organizations can detect, contain, and recover from incidents.

Q: How can organizations start preparing?

A: Organizations can stress-test their vulnerability management processes, modernize resilience infrastructure, and validate recovery capabilities. Acting during this early window provides a meaningful strategic advantage.

Tim Zonca is Vice President, Portfolio Management, at Commvault.

More related posts


Thumbnail_Blog-Anthropic-Project-ResOps-2026

Anthropic’s Project Glasswing Makes the Case for ResOps

Read more about Anthropic’s Project Glasswing Makes the Case for ResOps

Die wichtigsten Erkenntnisse

  • Agentische KI birgt neue Sicherheitsrisiken, da sie systemübergreifend plant, sich Dinge merkt und handelt, anstatt nach einem einzigen Prompt-Antwort-Zyklus aufzuhören.
  • Verfälschte Trainingsdaten können das Modellverhalten im großen Maßstab unbemerkt beeinflussen, selbst wenn das Modell bei Standardtests noch eine normale Leistung zu erbringen scheint.
  • Kompromittierte Vektordatenbanken können die Entscheidungen von Agenten beeinflussen, indem sie den Kontext verfälschen, auf den sich das Modell stützt, wodurch unerwünschtes Verhalten als legitim erscheint.
  • Die Identität unkontrollierter Agenten führt zu einem Problem bei der Zugriffskontrolle in Maschinen-Geschwindigkeit, für dessen Bewältigung herkömmliche, auf den Menschen ausgerichtete Identitätssysteme nicht ausgelegt sind.
  • Auf fehlerhaften Zuständen basierende, aufeinander aufbauende Entscheidungen können zu Fehlern in mehreren Agenten und Arbeitsabläufen führen, was ein Zurücksetzen und die Wiederherstellung erheblich erschwert.

The tools, controls, and governance policies most enterprises have in place were designed for systems that answer questions – retrieval tools, copilots, generative assistants. Systems that respond to a prompt and stop. When something went wrong, the failure was discrete. Fix the prompt, adjust the configuration, move on.

Agentic AI doesn’t work that way. These systems plan, remember, and execute across the enterprise without step-by-step human instruction. They maintain state. They coordinate with other agents. They act on production systems – writing to databases, triggering workflows, making decisions at machine speed.

That architectural shift introduces four threat vectors that existing security frameworks were never designed to address. If your AI governance strategy doesn’t account for them, you likely have exposure you probably can’t see.

1. Verfälschte Trainingsdaten

An AI system is only as trustworthy as the data it was trained on. That statement has always been true. What’s changed is the attack surface.

In agentic AI deployments, training pipelines are larger, more complex, and frequently assembled from multiple sources – internal data, third-party feeds, vendor-provided datasets. Each dependency in that chain is a potential injection point. An adversarial actor who can influence training data – through supply chain compromise, insider access, or contamination of a shared data source – can shape model behavior at scale.

What makes this particularly dangerous is that poisoned models often perform normally on standard benchmarks. The manipulation may be surgical: designed to produce specific outputs in specific contexts while behaving correctly everywhere else.

By the time the effect surfaces in production, the model has been in use for weeks or months, and tracing the contamination back to its source requires exactly the kind of relational data provenance most organizations don’t have.

The question to ask: Can you produce a complete, verifiable record of what data your models were trained on – at a specific point in time?

2. Kompromittierte Vektordatenbanken

Vector databases are the memory layer of agentic systems. Before an agent acts, it queries a vector store to retrieve relevant context – past interactions, domain knowledge, reference data – that shapes what it does next.

Most security teams aren’t thinking about vector databases the way they think about other sensitive data stores. They should be.

A compromised vector database doesn’t just return wrong answers. It shapes the decisions that follow. Injected embeddings – malicious content inserted into the vector store – can redirect agent behavior in ways that appear completely legitimate from the outside.

An agent asked to approve a transaction retrieves context that subtly reframes the approval criteria. An agent managing customer communications pulls context that steers responses in an attacker’s preferred direction. The action looks correct. The reasoning looks sound. But the underlying context has been manipulated.

This attack vector is particularly hard to detect because it operates below the model layer. Standard model monitoring won’t catch it. The model is behaving exactly as trained – it’s the context it’s reasoning from that’s been corrupted.

The question to ask: Is your vector database treated as a sensitive, governed data asset – with access controls, integrity monitoring, and audit logging comparable to your most critical production databases?

3. Identität eines nicht regulierten Akteurs

In a multi-agent architecture, agents don’t just interact with data – they interact with each other. They spawn subagents, delegate tasks, request outputs, and synthesize results from agents they’ve never been explicitly connected to. To do this, they authenticate, present credentials, and establish trust.

Agent identity is the access control layer for the autonomous enterprise – and it’s a gap that identity security vendors and identity providers (IDPs) don’t close. Their governance frameworks are built for human identity.

Agent identities created within those same rules appear completely legitimate: They were provisioned correctly, they followed policy. The IDP isn’t failing – it simply has no framework for determining whether an agent is acting outside the context it was created for, has been quietly escalated, or is coordinating where it shouldn’t be.

The exposure is qualitatively different from traditional credential compromise. When a human user’s credentials are stolen, the attacker operates within that user’s permissions, at human speed.

When an agent’s identity is compromised, the attacker gains access to the autonomous decision-making layer – the ability to trigger workflows, approve actions, coordinate with other agents, and exfiltrate data at machine speed, at scale, through channels that appear entirely normal.

Identity-layer failures are also among the hardest to detect after the fact. Agent actions taken under a compromised identity don’t look anomalous – they look like legitimate agent behavior. And because they’re generated by a system rather than a human, the volume can be enormous before anyone notices.

Recovery compounds the problem. Most AI recovery playbooks focus on restoring data: training sets, model weights, pipeline configurations. Identity is rarely on the list. A system recovered with clean data but misaligned identity configurations isn’t actually recovered. It’s a clean system with a poisoned access layer.

The question to ask: Is agent identity managed with the same rigor as human identity – with lifecycle management, least-privilege access, and inclusion in recovery playbooks?

4. Kaskadierende Entscheidungen, die auf einer falschen Zustandsangabe beruhen

The first three attack vectors are discrete. This one is systemic – and in many ways it can be the most difficult to contain.

Multi-agent architectures are designed for coordination. Agents share context, pass outputs to one another, and build on each other’s work. That coordination is what makes them powerful. It’s also what makes failures propagate.

An agent operating on corrupted memory doesn’t fail cleanly. It produces outputs – decisions, actions, data – that other agents consume. Those agents produce their own outputs. By the time the original corruption surfaces as something observable, bad state may have touched dozens of downstream processes, across multiple agents, with no clean rollback path.

This is what makes the context gap so significant. At any given moment, your AI system consists of a model version, a set of training data, an artifact store, a pipeline configuration, and a set of active agent interactions – all of which need to reflect the same operational state to constitute a trustworthy, recoverable system. When they don’t, you don’t just have an error. You have a system that is coherent in pieces and incoherent as a whole.

Point tools can each confirm their own slice. None can confirm the pieces belong together. That’s not a monitoring problem you can solve by adding another tool. It’s a structural gap – and the only way to close it is with a system that captures AI state relationally: what was running, against what data, with what configuration, at what moment.

The question to ask: If your AI infrastructure were compromised today, could you identify exactly what state every component was in before the incident – and prove it?

Was dies für Ihre Sicherheitsstrategie bedeutet

Each of these four vectors requires a different defensive response. But they share a common implication: The governance and resilience frameworks designed for the previous era of AI don’t cover the failure modes of the agentic era.

Securing agentic AI requires extending your framework in three directions:

  • Deeper, into the data and identity layers that sit below the model.
  • Broader, to cover agent-to-agent interactions that existing monitoring doesn’t observe.
  • Relationally, to capture not just the state of individual components, but how they fit together at any point in time.

That last requirement is the one most organizations haven’t yet confronted. And it’s the one that will determine whether, when something goes wrong, you have a recoverable system or a collection of accurate-looking reports describing something that no longer exists.

Read „The Agentic Blind Spot: Why AI Resilience Demands a System of Record“, um zu erfahren, warum Sie ein SOR benötigen, um die Konsistenz und Genauigkeit Ihrer KI-Daten zu gewährleisten.

FAQs

Q: Why are agentic AI systems riskier than traditional generative AI tools?

A: Agentic systems do more than answer prompts. They maintain state, coordinate with other agents, and take actions in production environments, which expands the attack surface far beyond simple prompt manipulation.

Q: What makes poisoned training data so difficult to detect?

A: The manipulation can be highly targeted, affecting only specific situations while leaving normal benchmarks intact. That means a model may look healthy until the poisoned behavior appears in real use.

Q: How can a vector database become a security problem?

A: A vector database shapes the context an agent uses before acting. If that context is altered, the agent may make decisions that seem reasonable on the surface but are really being guided by malicious data.

Q: Why is agent identity different from human identity?

A: Agent identity is tied to autonomous actions, delegation, and machine-speed execution. Traditional identity governance is designed for people, so it often misses whether an agent is acting outside its intended context.

Q: Why is cascading bad state such a serious issue in multi-agent systems?

A: Once one agent consumes corrupted output, that error can spread to downstream agents and workflows. The result is not just one bad decision, but a chain of connected failures.

Q: How can organizations improve AI security?

A: Extend governance deeper into data and identity layers, monitor agent-to-agent interactions, and track AI state relationally to enable them to reconstruct what happened during an incident.

Michael Thelander is Senior Director, Product Marketing, at Commvault.

Verwandte Blogs

More related posts


Thumbnail_Blog-Data-Access-Governance-2026

Securing AI with Unified Data Access Governance

Read more about Securing AI with Unified Data Access Governance
Thumbnail_Blog-Environmental-Footprint-AI-2026

Smarter Data, Greener AI

Read more about Smarter Data, Greener AI
Thumbnail_Blog-Anthropic-Project-ResOps-2026

Anthropic’s Project Glasswing Makes the Case for ResOps

Read more about Anthropic’s Project Glasswing Makes the Case for ResOps
Thumbnail_Blog-Data-Rooms-2025-Linkedin

Data Activate: Unlocking the Power of Trusted Data for AI Innovation

Read more about Data Activate: Unlocking the Power of Trusted Data for AI Innovation
Thumbnail_Blog-AI-Agents-2026

AI Agents Are Everywhere. Do You Know What They’re Doing?

Read more about AI Agents Are Everywhere. Do You Know What They’re Doing?
Thumbnail_Blog-Building-AI-Agents-2026

From Experimentation to Operation: Building AI Agents You Can Actually Trust

Read more about From Experimentation to Operation: Building AI Agents You Can Actually Trust

Die wichtigsten Erkenntnisse

  • Agentische KI-Systeme sind zustandsbehaftet und arbeiten kontinuierlich, weshalb herkömmliche Wiederherstellungsmodelle nicht ausreichen.
  • Die Speicherebene (Vektordatenbanken und Kontextspeicher) ist eine kritische, aber nur unzureichend überwachte Angriffsfläche.
  • Workflows zur Entscheidungsfindung zur Laufzeit können angepasst werden, ohne dass dabei herkömmliche Sicherheitswarnungen ausgelöst werden.
  • Aufgrund von Lücken in der Beobachtbarkeit bei Interaktionen zwischen Agenten haben die meisten Unternehmen nur einen unvollständigen Überblick über die Risiken.
  • Eine echte Wiederherstellung erfordert eine einheitliche, zeitlich abgestimmte Aufzeichnung aller Systemebenen, um einen vertrauenswürdigen Zustand wiederherzustellen.

Most enterprises entering the agentic AI era are managing resilience with the wrong mental model – and the data backs it up: Nur jedes fünfte Unternehmen verfügt über ein ausgereiftes Modell zur Steuerung autonomer KI-Agenten. They’re thinking about AI the way they think about applications: discrete, stateless, recoverable by restoring clean data to a clean environment.

Agentic AI doesn’t work that way. These systems are stateful, continuously operating, and architecturally layered in ways that create failure modes most security and resilience frameworks weren’t designed to address. The gap isn’t in tooling. It’s in understanding what’s actually running – and what “recovery” has to mean for systems built this way.

There are four architectural layers that define the problem. Each one is distinct. Each one is underprotected. And together, they explain why an agentic AI system can appear recoverable while remaining fundamentally compromised.

Layer 1: Agent Memory – The Attack Surface You’re Not Watching

Traditional enterprise applications don’t remember anything between sessions. Agentic AI does. The memory layer – primarily vector databases storing embeddings, but also session state and retrieved context – is what gives agents continuity across interactions. It’s what allows an agent to pick up where it left off, to draw on prior context, to build a coherent picture of a complex workflow over time.

It is also one of the most consequential attack surfaces in the modern enterprise stack – and one of the least monitored.

The attack vector is subtle enough to evade most conventional security tooling. An adversary who can influence what gets written to a vector database can shape what the agent believes to be true. Injected or manipulated embeddings don’t need to look malicious – they need to look authoritative.

A compromised memory store can redirect agent behavior, exfiltrate data through agent actions, or cause an agent to make decisions that appear legitimate but serve an attacker’s objectives. None of this requires touching the model itself.

The detection problem is compounded by the volume and velocity of vector database writes in active agentic deployments. Anomaly detection tools built for structured data don’t translate well to embedding space. The signal is there – but most organizations aren’t equipped to read it.

What resilience requires here: continuous integrity monitoring of vector databases, not just backup. Version-controlled embeddings with a provable chain of custody. The ability to identify, at any point in time, exactly what the memory layer contained – and to restore to a verified clean state, not just a recent one.

Layer 2: Runtime Control – When the Workflow Is the Threat

Agentic AI doesn’t execute fixed scripts. It plans. At runtime, an agent receives a goal, determines the steps required to achieve it, selects the tools it needs, and executes – often spawning subagents to handle parallel workstreams. The workflow is dynamic, constructed in the moment, and frequently long-running.

This is what makes agentic AI genuinely useful. It’s also what makes it genuinely difficult to protect.

In a conventional automation environment, a compromised workflow is bounded. It does what it was configured to do, and it stops. A compromised agentic workflow is different: It adapts.

If an attacker can influence the planning layer – through a poisoned prompt, a manipulated tool response, or a corrupted planning model – the agent will pursue the attacker’s objective using whatever legitimate tools and access it has. It will look like normal operation. The logs, to the extent they exist, will show authorized tool calls.

Consider a procurement agent tasked with validating vendor invoices against contract terms. Under normal operation, it checks invoice amounts, cross-references approval thresholds, and flags exceptions for human review.

An attacker who can influence the planning layer – through a manipulated tool response from the contract database – doesn’t need to touch the approval logic directly. They simply give the agent a contract record with altered thresholds.

The agent plans correctly against corrupted inputs. Every tool call it makes is legitimate. Every decision it reaches is wrong. By the time the anomaly surfaces in a finance reconciliation, the workflow has processed weeks of invoices and the audit trail shows nothing but authorized actions.

The window between compromise and detection in these scenarios is not measured in seconds. Agentic workflows operate continuously. By the time anomalous outcomes surface, the workflow may have touched dozens of systems, made hundreds of decisions, and left changes across production environments that are difficult to enumerate and harder to reverse.

What resilience requires here: runtime monitoring that watches what agents are deciding, not just what they’re doing. Intervention mechanisms that can halt a running workflow cleanly without cascading failures. Recovery playbooks built for long-running agentic processes – not just for discrete transactions.

Layer 3: Agentic Observability – The Logging Gap at Machine Speed

Enterprise logging infrastructure was built for human-scale operations. It captures what systems do, at a granularity and latency designed for human review. Agentic AI operates at a different speed entirely.

In an active multi-agent deployment, agents are spawning subagents, passing context between one another, making tool calls, and synthesizing outputs – continuously, in parallel, faster than conventional logging pipelines were designed to capture.

The interactions that matter most for security – agent-to-agent communications, context handoffs, tool invocations that cross trust boundaries – are exactly the interactions that existing monitoring frameworks leave most underobserved.

Today, nur 17 % der Unternehmen continuously monitor agent-to-agent interactions. The other 83% are governing agentic AI based on a partial picture – one that captures what individual agents do in isolation but misses the interaction layer where the most consequential security events occur.

This isn’t a gap that more logging volume solves. The problem isn’t the quantity of data being captured – it’s that the data structures and latency requirements of agentic interactions don’t fit well into observability frameworks designed for slower, more structured systems. Closing this gap requires purpose-built agentic observability tooling, or significant adaptation of existing infrastructure.

What resilience requires here: end-to-end visibility into agent-to-agent interactions, not just individual agent outputs. Logging architectures that can operate at agentic speed without dropping events. The ability to reconstruct, after the fact, the full sequence of agent decisions and interactions for any given workflow.

Layer 4: Multi-Agent Coordination – Where Emergent Failures Hide

The most architecturally novel risk in agentic AI doesn’t come from any single compromised agent. It comes from how agents depend on one another – and how failures propagate across those dependencies before anyone realizes something is wrong.

In a multi-agent architecture, agents share context. An orchestrator agent passes a task brief to a subagent; the subagent returns a result that the orchestrator incorporates into its next decision.

If the subagent’s output is corrupted – through a compromised memory layer, a manipulated tool response, or a poisoned planning model – the orchestrator has no native way to detect it. It treats the output as authoritative. It incorporates it. It acts on it. And it passes its own now-compromised output downstream.

This is the emergent failure mode: a corruption that originates in one layer, propagates through agent interactions, and surfaces as an anomalous outcome in a system several steps removed from the original compromise. By the time it’s visible, the causal chain is long and the blast radius is significant.

Consider a threat intelligence pipeline where a data-gathering agent ingests feeds from external sources, a classification agent categorizes and scores them, and an orchestrator incorporates the scored intelligence into security posture recommendations pushed to downstream teams.

If the data-gathering agent’s memory layer is compromised – subtly, through injected embeddings that cause it to weight certain threat actors as low-risk – the classification agent receives inputs it has no reason to question. It classifies accurately against what it’s given.

The orchestrator incorporates the results confidently. Security teams downstream deprioritize the relevant threat category based on what looks like a coherent, multi-source consensus. The failure originated in Layer 1. It expressed itself in Layer 4. Nothing in between flagged an anomaly because nothing in between had visibility across the full chain.

The governance frameworks most enterprises apply to AI were designed for model outputs – what the AI says. Multi-agent coordination failures are not model output failures. They are systems failures, arising from the interaction layer between models, and they require a different kind of governance: one that monitors and controls not just individual agent behavior but the trust relationships between agents, the integrity of context as it passes between them, and the access rights that govern what any agent can request of any other.

What resilience requires here: agent identity management that treats inter-agent trust as a first-class security concern. Integrity verification for context as it moves across agent boundaries. Governance policies that cover autonomous agent behavior – not just the outputs of individual models.

Das Beziehungsproblem, das alle vier miteinander verbindet

These four layers are distinct in their failure modes, but they share a common vulnerability: none of them has a shared record of how they relate to each other at a specific point in time.

The model registry knows what version is running. The vector database knows what’s in memory. The orchestration layer knows what workflow is active. The identity system knows what agents have what access. Each can confirm its own slice of the picture. None can confirm whether those slices belong together – whether they reflect the same operational state, the same moment, the same trustworthy configuration.

That’s the context gap. And it’s why recovery from an agentic AI compromise isn’t a data restoration problem. It’s a coherence problem – one that requires a unified record of the relationships between layers, not just the components themselves.

Close this gap before an incident, or spend an incident trying to close it.

The architecture challenges covered here are only part of what security and resilience leaders need to understand about agentic AI risk. „The Agentic Blind Spot: Why AI Resilience Demands a System of Record“ goes further, examining where most enterprises actually stand on AI resilience readiness, what the governance gaps look like in practice, and what it takes to make “our AI is trustworthy” a provable claim, not just an assertion.

FAQs

Q: Why doesn’t traditional disaster recovery work for agentic AI?

A: Traditional recovery assumes systems are stateless and can be restored from clean backups. Agentic AI systems retain memory, evolve over time, and depend on layered interactions, making simple restoration insufficient to regain trust.

Q: What makes the memory layer in agentic AI vulnerable?

A: The memory layer stores embeddings and contextual data that influence agent decisions. If compromised, attackers can subtly manipulate what the agent “believes,” leading to incorrect but seemingly legitimate actions.

Q: How can attackers exploit runtime workflows in agentic AI?

A: Attackers can influence planning inputs, prompts, or tool responses, causing agents to execute harmful actions using legitimate processes. These actions often appear normal in logs, making detection difficult.

Q: Why is observability a challenge in multi-agent systems?

A: Agentic systems operate at machine speed with continuous interactions between agents. Traditional logging systems are not designed to capture or process this level of dynamic, high-frequency activity.

Q: What are emergent failures in multi-agent environments?

A: Emergent failures occur when a small compromise in one agent or layer propagates across interconnected agents, resulting in large-scale issues that are difficult to trace back to the original source.

Q: What does effective recovery look like for agentic AI?

A: Effective recovery requires more than restoring data – it demands a coherent snapshot of all system layers, including memory, workflows, identities, and interactions, aligned to a verified trustworthy state.

Tim Zonca is Vice President, Portfolio Management, at Commvault.

More related posts


Thumbnail_Blog-Architect-for-tomorrow-2026

Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience

Read more about Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience
Thumbnail_Blog-Dangerous-Silos-IDC-Resops-2026 (1)

The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan

Read more about The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan
Thumbnail_Blog-IDC-Resops-2026

Business Continuity Planning for the Cloud-Native Era

Read more about Business Continuity Planning for the Cloud-Native Era

Scaling a data-driven company is hard. Scaling one while meeting GDPR requirements, managing thousands of customers, enabling analytics teams, and standing up new infrastructure in under two weeks? That’s a different level of complexity.
In a recent episode of STRIVE, I sat down with Asif Dromi of monday.com and Ben Herzberg of Commvault to unpack what it really takes to operationalize data security at scale – not in theory, but in practice. This isn’t a high-level conversation about best practices. It’s a real-world look at how security, compliance, automation, and infrastructure decisions intersect when the clock is ticking.
Watch the gesamte Folge.
If you’re a CISO, data leader, architect, or compliance owner, this episode gives you something more valuable than theory. It shows how:

  • Ein schnell wachsendes Unternehmen bewältigte die Herausforderungen der DSGVO, ohne dabei die Innovation zu bremsen.
  • „Infrastructure as Code“ kann Audits vereinfachen.
  • Automatisierung verringert Risiken, anstatt die Komplexität zu erhöhen.
  • Security and business agility don’t have to compete.

It’s rare to hear directly from operators who’ve done this under real constraints. That’s what makes this STRIVE conversation different.

Die wichtigsten Erkenntnisse: Umsetzung von Datensicherheit in großem Maßstab

  • Compliance and growth don’t have to compete. Monday.com demonstrates how GDPR requirements and rapid expansion can coexist when security is built into architecture from the start.
  • Manual permissions don’t scale. Automation does. Infrastructure as code and API-driven access controls can turn governance from a bottleneck into a force multiplier.
  • Role-based access must evolve with data usage. As more teams depend on analytics, visibility and fine-grained controls become important to help prevent permission sprawl.
  • Operationalized security means visibility. It’s not just about setting policies – it’s about monitoring, auditing, and adapting controls dynamically as environments change.
  • Speed is possible when architecture is intentional. A compliant European data warehouse stood up in under two weeks because governance, automation, and tooling were designed to scale.
  • Security maturity enables innovation. When permissions, infrastructure, and compliance are programmable, organizations can move faster.

Die eigentliche Herausforderung: Wachstum + Compliance + Geschwindigkeit

For monday.com, the challenge wasn’t just storing European data in Europe. It was:

  • Gewährleistung der DSGVO-Konformität und der regionalen Datenspeicherung.
  • Sicherstellen, dass die Mitarbeiter nur auf relevante Daten zugreifen.
  • Gewährleistung von Transparenz und Nachvollziehbarkeit.
  • Unterstützung von Analysten und Entwicklern, die schnellen Zugriff benötigten.
  • Und das alles unter hohem Zeitdruck.

As Asif explains in the episode, becoming a data-driven organization means internal access expands rapidly. The more teams rely on analytics, the more complex permissions become.
And that’s where many organizations hit a wall. Security becomes manual, permissions become fragile, and compliance becomes reactive. That’s not operationalized security. That’s a house of cards.

Designing Security into the Architecture from Day One 

Einer der spannendsten Aspekte dieser Folge ist, wie monday.com das Problem aus architektonischer Sicht angegangen ist. Anstatt die Compliance nachträglich einzubauen, hat das Unternehmen Folgendes entwickelt:

  • Ein eigens dafür vorgesehenes europäisches Data Warehouse.
  • Klare, rollenbasierte Zugriffskontrollen.
  • Feinkörnige Berechtigungsmodelle.
  • Automatisierte Governance-Ebenen.

Ben beschreibt, was in vielen großen Unternehmen geschieht: Im Laufe der Zeit häufen sich Berechtigungen in mehreren Schichten an, oft ohne zentrale Übersicht. Schließlich weiß niemand mehr genau, wer auf welche Daten zugreifen darf. Sicherheit zu operationalisieren bedeutet, diese Fehlentwicklung zu vermeiden. Es bedeutet, Systeme zu entwickeln, bei denen sich die Governance automatisch mit steigender Nutzung skaliert.

Automatisierung ist der Kraftmultiplikator

If there’s one theme that runs through this episode, it’s automation. Instead of treating permissions as tickets and manual updates, monday.com wrapped their infrastructure in code. Databases, roles, and access policies could be created and modified programmatically.
The result? A compliant, scalable environment stood up in less than two weeks. That’s not luck. That’s architecture. And it’s a powerful reminder that security doesn’t slow you down when it’s built correctly. It enables speed.

Was die Umsetzung von Datensicherheit wirklich bedeutet

“Operationalizing” gets used a lot. In this episode, it’s defined as:

  • Kontinuierliche Transparenz bei sensiblen Daten.
  • Zentralisierte und automatisierte Berechtigungsverwaltung.
  • Zugriffsüberwachung.
  • Integration mit Tools für die Zusammenarbeit.
  • Richtlinien, die sich an die steigende Nutzer- und Datenmenge anpassen.

Static controls don’t scale. Manual workflows don’t scale. Security must become dynamic – part of the operating fabric of the organization. And that shift is where many enterprises struggle today.

Die komplette Folge von STRIVE ansehen

In the discussion, you’ll hear more about:

  • Wie monday.com sein europäisches Data Warehouse aufgebaut hat.
  • Die wichtigsten Erkenntnisse aus der raschen Umsetzung.
  • Warum Automatisierung unverzichtbar war.
  • Was Unternehmen oft unterschätzen, wenn es um die Ausuferung von Berechtigungen geht.
  • Wie man die Umsetzung von Governance-Maßnahmen angeht, bevor KI-Initiativen ausgeweitet werden.

Jetzt anschauen.

FAQs 

Q: How can small teams implement scalable data security?

A: Start with a clear permissions model and infrastructure-as-code tools. Automate permission management early to help avoid manual bottlenecks as you grow.

Q: What role does automation play in compliance?

A: Automation helps enable consistency, reduce errors, and simplify audits. Using APIs and scripts, you can monitor and adjust permissions dynamically.

Q: How long does it typically take to set up a compliant, scalable data environment?

A: With the right planning and tools, organizations like monday.com have achieved this in less than two weeks. Speed depends on scope and existing infrastructure.

Q: What are best practices for operationalizing data security?

A: Implement role-based access controls, automate permission management, monitor access logs regularly, and integrate security tools with collaboration platforms for real-time oversight.

Chris Mierzwa is Senior Director, Portfolio Marketing, at Commvault.

More related posts


Thumbnail_Blog-GoogleWorkspace-2026

Expanding Google Workspace Protection with Commvault eDiscovery

Read more about Expanding Google Workspace Protection with Commvault eDiscovery
Thumbnail_Blog-Data-Leakage-Loops-2026

Are You Ready for Data Leakage Loops?

Read more about Are You Ready for Data Leakage Loops?
Thumbnail_Blog-Tornado-2025-Linkedin

The Trust Tightrope: Why New Yorkers Demand More from Businesses Than They Do from Themselves

Read more about The Trust Tightrope: Why New Yorkers Demand More from Businesses Than They Do from Themselves
Thumbnail_Blog_FinServ-Cybersecurity-2025

Modernizing Financial Cybersecurity: From Reactive to Resilient

Read more about Modernizing Financial Cybersecurity: From Reactive to Resilient

Die wichtigsten Erkenntnisse

  • Compliance-Rahmenwerke halten die aus tatsächlichen Fehlschlägen gewonnenen Erkenntnisse fest und helfen Organisationen dabei, ihre Widerstandsfähigkeit, Governance und operative Stabilität zu stärken.
  • Unternehmen, die Compliance als Maßnahme zum Aufbau von Vertrauen betrachten, können dazu beitragen, das Vertrauen der Kunden, die Beziehungen zu Partnern und die Glaubwürdigkeit der Marke zu stärken.
  • Die Angleichung der regulatorischen Rahmenbedingungen und strenge Risikokontrollen können dazu beitragen, die Versicherungsergebnisse zu verbessern, indem sie eine ausgereifte und widerstandsfähige Sicherheitslage demonstrieren.
  • Durch die Verknüpfung von Compliance-Anforderungen mit messbaren Geschäftsergebnissen können Unternehmen ihre Investitionen in die Widerstandsfähigkeit direkt mit dem Schutz ihrer Einnahmen und der Geschäftskontinuität in Verbindung bringen.
  • Maßnahmen zur Stärkung der Cyber-Resilienz wie unveränderliche Backups, schnelle Wiederherstellung und Governance-Rahmenwerke helfen Unternehmen dabei, Compliance in einen Wettbewerbsvorteil zu verwandeln.

In boardrooms across Europe and beyond, compliance has become a loaded word. It conjures images of endless documentation, mounting regulatory pressure, and the looming threat of fines.
GDPR. NIS2. DORA. The acronyms keep coming, and for many organizations, it can feel like they are choking on regulation.
But what if we’ve been looking at compliance the wrong way? What if compliance isn’t just about avoiding penalties – but about building a better, stronger, more resilient business?

Die Versicherungsanalogie: Regeln, die aus gutem Grund existieren

Thier’s a useful parallel between compliance and insurance.
When you insure your car, the insurer sets certain conditions. Your brakes must work. Your tires shouldn’t be bald. An alarm system might be required. You can argue about the inconvenience, or the cost – but fundamentally, those rules exist because they help reduce risk. They help make accidents less likely. They help protect both you and others.
And hier’s the key point: Those requirements are usually a good idea, whether you buy the insurance or not.
Regulation works in much the same way. Governments and regulators don’t create frameworks because they enjoy it. Regulations are responses to real-world failures – data breaches, operational disruptions, systemic risk. They codify lessons learned the hard way.
You may object to the burden. You may find it frustrating. But when you look closely at what these frameworks require, it’s hard to argue that the core principles are unsound.

  • Schützen Sie Kundendaten.
  • Sicherstellung der betrieblichen Ausfallsicherheit.
  • Machen Sie sich mit den Risiken Ihrer Lieferkette vertraut.
  • In der Lage sein, sich von Cybervorfällen zu erholen.
  • Zeigen Sie, dass Sie verantwortungsbewusst handeln und Rechenschaft ablegen.

Das ist alles keine schlechte Idee.

Von der Vermeidung von Bußgeldern bis hin zur Schaffung von Vertrauen

Too often, compliance is framed defensively: “Do this so you don’t get fined.” “Do this so you don’t go to jail.”

That’s a low bar. And it’s a missed opportunity. When we shift the perspective, compliance becomes something much more powerful. It becomes a driver of trust.
Take GDPR as an example. At its heart, it’s about protecting personal data. If your organization implements strong data protection practices – not just to tick a box, but because your systems genuinely safeguard customer information – that builds trust. Customers are more confident doing business with you. Partners are more willing to integrate with you. Regulators view you as lower risk.
Trust is not a regulatory outcome. It’s a commercial advantage.
The same applies to the Digital Operational Resilience Act. It’s not just about reporting incidents; it’s about being able to withstand and recover from disruption. In a world whier cyberattacks are inevitable, resilience is not optional. It’s foundational to continuity, reputation, and long-term value.
When compliance drives resilience, resilience drives business stability – and stability drives growth.

Regulierung und Versicherung: Ein Regelkreis

Thier’s also a natural alignment between regulation and insurance markets. When regulators mandate certain standards, insurers quickly follow. Organizations that demonstrate compliance and strong risk controls are more attractive to underwriters. They may benefit from better terms, broader coverage, or more favorable premiums.
This creates a reinforcing cycle:

  • Die Verordnung legt Mindeststandards fest.
  • Unternehmen verstärken ihre Kontrollmaßnahmen.
  • Versichierr belohnen ein stärkeres Risikobewusstsein.
  • Die Märkte werden stabiler und widerstandsfähiger.

Compliance ist in diesem Zusammenhang ein Signal an den Markt: Wir nehmen Risiken ernst.

Das fehlende Glied: Die Verknüpfung von Compliance mit Geschäftsergebnissen

One of the most important opportunities for organizations – particularly technology providers – is to make the “line of sight” between compliance and business value explicit.
For example:

  • Wenn ein Produkt unveränderliche Sicherungskopien erstellt, trägt dies zur Einhaltung gesetzlicher Anforderungen hinsichtlich der Datenintegrität bei.
  • Wenn dies eine schnelle Wiederherstellung nach Cybervorfällen ermöglicht, trägt dies zur Erfüllung der Vorgaben zur operativen Widerstandsfähigkeit bei.
  • Wenn es klare Prüfpfade und Berichtsfunktionen bietet, trägt dies zur Erfüllung der Anforderungen an Governance und Aufsicht bei.

But it shouldn’t stop thier. The next step is to articulate the business benefit:

  • Immutable backups help reduce the impact of ransomware – and protect revenue.
  • Faster recovery helps minimize downtime – and preserves customer confidence.
  • Strong governance helps reduce regulatory scrutiny – and enhances brand credibility.

This mapping is critical. Compliance is not the end goal; it’s the mechanism that enables the outcomes that businesses care about: continuity, reputation, customer trust, and competitive differentiation.

Compliance als Innovation, nicht als Verpflichtung

Thier’s a tendency to treat compliance as a “get-it-done” exercise. A cost center. A necessary evil.
But if we look at history, many best practices that are now considered fundamental to modern IT and security originated in regulatory or insurance requirements. Over time, they became embedded in how well-run organizations operate.
Encryption. Access controls. Incident response planning. Business continuity testing. Third-party risk management.
At one time, these may have been viewed as regulatory burdens. Today, they are table stakes for any serious enterprise.
The organizations that treat compliance as an innovation catalyst – rather than a checkbox exercise – are often the ones that pull ahead. They embed resilience into their architecture. They design with governance in mind. They turn regulatory requirements into product capabilities and customer value propositions.

Cyber-Resilienz: Wo Compliance und Strategie aufeinandertreffen

This is whier cyber resilience becomes central.
Modern regulations increasingly recognize a simple truth: Prevention is not enough. Incidents will happen. The differentiator is how well an organization can respond and recover.
Cyber resilience – the ability to withstand, recover from, and adapt to cyber disruption – is no longer just a security concern. It’s a strategic imperative. It supports regulatory compliance, yes. But more importantly, it underpins operational continuity and business confidence.
When organizations invest in resilient architectures, immutable data, rapid recovery capabilities, and robust governance frameworks, they are not merely satisfying regulators. They are building durable enterprises.

Ein anderer Blick auf Compliance

Perhaps it’s time to change the narrative.
Instead of asking, “What’s the minimum we need to do to comply?” we should be asking:

  • Inwiefern macht uns diese Verordnung stärker?
  • Welche bewährte Praxis wird hier festgeschrieben?
  • Wie können wir dies nutzen, um das Vertrauen unserer Kunden und Partner zu stärken?
  • Inwiefern verschafft dies einen Wettbewerbsvorteil?

Compliance done well is not about fear. It’s about foresight.
It reflects lessons learned across industries. It embeds best practice into everyday operations. And when connected clearly to product capabilities and business outcomes, it becomes a powerful commercial story.
Yes, regulation can feel burdensome. Yes, the acronyms keep coming. But underneath the paperwork lies something far more valuable: a framework for running a better business.
Compliance isn’t just about avoiding penalties. It’s about enabling resilience. And resilience, ultimately, is what drives sustainable success. Learn more about how Commvault enables data protection to help your organization meet compliance requirements hier.

FAQs

Q: Why should organizations view compliance as more than a regulatory obligation?

A: Compliance frameworks often reflect best practices developed in response to real-world cyber incidents, operational failures, and governance challenges. Organizations that embrace compliance strategically can help strengthen resilience, improve trust, and create long-term business value.

Q: How does compliance contribute to customer trust?

A: Strong compliance practices demonstrate that an organization takes data protection, governance, and operational continuity seriously. This can help increase customer confidence, strengthen partner relationships, and position the organization as a lower-risk business.

Q: What is the connection between compliance and cyber resilience?

A: Modern regulations increasingly focus on an organization’s ability to recover from disruptions rather than solely preventing them. Investments in resilient infrastructure, immutable backups, and rapid recovery capabilities can help organizations maintain continuity during cyber incidents.

Q: How can compliance positively impact insurance and risk management?

A: Organizations with mature compliance programs and strong security controls are often viewed more favorably by insurers. This can lead to better coverage options, improved policy terms, and potentially lower premiums.

Q: Why is it important to connect compliance initiatives to business outcomes?

A: Compliance efforts are most effective when organizations clearly demonstrate how controls support broader goals such as protecting revenue, reducing downtime, and preserving customer trust. This helps leadership view compliance as a strategic investment rather than a cost center.

Q6: How can organizations turn compliance into a competitive advantage?

A: Businesses that embed resilience, governance, and security into their products and operations can differentiate themselves in the market. By proactively aligning with regulatory expectations, organizations can strengthen their reputation and create greater confidence among customers and stakeholders.

Darren Thomson is Field CTO at Commvault.

More related posts


Thumbnail_Blog-Architect-for-tomorrow-2026

Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience

Read more about Architect for Tomorrow: Unified Data Protection as the Foundation for Resilience
Thumbnail_Blog-Dangerous-Silos-IDC-Resops-2026 (1)

The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan

Read more about The Importance of Recovery Point Objective (RPO) in Your Business Continuity Plan
Thumbnail_Blog-IDC-Resops-2026

Business Continuity Planning for the Cloud-Native Era

Read more about Business Continuity Planning for the Cloud-Native Era

Die wichtigsten Erkenntnisse

  • Bei der operativen Souveränität geht es darum, wer Zugriff auf Systeme hat und unter welcher Rechtshoheit diese betrieben werden.
  • Der Zugriff von Anbietern, Telemetrie-Datenströme und Support-Prozesse können zu versteckten Souveränitätslücken führen.
  • Die operative Souveränität lässt sich schwerer nachweisen, da sie eine kontinuierliche Überwachung und Prüfung erfordert.
  • Unternehmen müssen in der Lage sein, jeden Zugriffsweg in geschützte Umgebungen nachzuweisen und zu dokumentieren.

Ask most organizations where their sovereignty program is strongest, and the answer is usually some version of the same two things: data locality and encryption. They know where their primary data lives. They’ve implemented bring-your-own-key or hold-your-own-key arrangements. They can point to certifications.

Ask them who accessed their sovereign environment in the last ninety days, from which countries, and under which legal jurisdictions – and the confidence tends to evaporate.

Operational sovereignty is the hardest pillar to audit, the most likely to be underestimated, and the most common place where a sovereignty posture that looks solid on paper breaks down in practice. The „„Digital Sovereignty Readiness Report““ names it as one of the four pillars – this post goes further.

The question most organizations can’t answer: ‘Who accessed your sovereign environment in the last 90 days, from which countries, and under which legal jurisdictions?’

Was operative Souveränität eigentlich bedeutet

Operational sovereignty is not about where data lives. It’s about who runs the environment – and who can reach it. It covers three things that most sovereignty programs treat as implementation details rather than first-class concerns:

  • Personnel access and jurisdiction. Every person who can access your sovereign environment – for support, maintenance, monitoring, or incident response – operates under a defined legal jurisdiction. If a support engineer in a country subject to a foreign data access law can reach your systems, the sovereignty of your infrastructure is only as strong as that engineer’s legal exposure.

Most organizations, when they audit this for the first time, find at least one support pathway that crosses a jurisdiction boundary they hadn’t mapped.

  • Third-party and vendor access. Your sovereignty boundary extends to every vendor, managed service provider, and software platform with access to your sovereign environment. ITSM platforms, monitoring tools, SIEM systems – if these sit outside your sovereignty boundary but have access to data or metadata within it, you have a gap that data locality controls cannot close.
  • Telemetry, billing, and control-plane traffic. Data sovereignty programs focus on primary data. Operational sovereignty requires mapping where everything else goes: the telemetry your infrastructure generates, the metadata your monitoring systems collect, the billing data your provider processes. These flows can cross jurisdiction boundaries even when primary data doesn’t – and they are rarely mapped.

Why This Pillar Is Harder To Certify – and Why That Matters

Data locality is relatively straightforward to document. You can point to a storage region, a data residency agreement, a third-party audit. Operational sovereignty doesn’t have the same paper trail. There is no certification that guarantees the jurisdictional status of every support engineer who might access your environment.

This is precisely what makes it both the hardest pillar to audit and the most important to get right. It also connects directly to the minimum viable sovereignty challenge: applying the right operational controls to the right workloads requires knowing what those controls are – and operational sovereignty is where that knowledge is most commonly absent.

Die Dimension der Lieferkette

NIS2, das die Cybersicherheitsverpflichtungen auf die Sektoren Energie, Verkehr, Gesundheitswesen und digitale Infrastruktur ausweitet, verpflichtet Organisationen nun dazu, die Cybersicherheitspraktiken ihrer Technologieanbieter zu bewerten. Für Souveränitätsprogramme hat dies direkte Auswirkungen: Die Haltung der Anbieter in Bezug auf Souveränität ist nicht länger nur eine nette Geste bei der Beschaffung. Sie ist eine überprüfbare Anforderung.

Das bedeutet, dass Sie jedem Anbieter innerhalb Ihrer Souveränitätsgrenzen neue Fragen stellen müssen: Wo befindet sich Ihr Supportpersonal? Unter welcher rechtlichen Zuständigkeit arbeitet es? Was geschieht mit den Zugriffsrechten auf meine Umgebung, wenn Ihr Unternehmen von einem Nicht-EU-Unternehmen übernommen wird?

So sieht es gut aus

Ein operativ souveränes Umfeld weist vier Merkmale auf, die sich nicht nur dokumentieren, sondern auch nachweisen lassen:

  • Every access pathway into the sovereign environment is mapped – not just primary access, but vendor access, support access, and monitoring system access.
  • Der Zugriffsstatus jeder Person oder jedes Systems mit diesem Zugriff wird dokumentiert und in festgelegten Abständen überprüft.
  • Telemetrie-, Metadaten- und Control-Plane-Datenverkehrsströme werden erfasst und entweder innerhalb der Hoheitsgrenze gehalten oder ausdrücklich geprüft und als außerhalb des Geltungsbereichs liegend akzeptiert.
  • The organization can answer the ninety-day access question – precisely, with evidence.

One more thing: Operational sovereignty doesn’t end at access control. If recovery requires personnel who operate outside your sovereignty boundary, the posture fails at the moment of an incident. That’s the subject of the fourth post in this series. Der „„„Digital Sovereignty Readiness Report“““ enthält eine direkte Bewertungsfrage zur operativen Souveränität.

FAQs

Q: What is operational sovereignty?

A: Operational sovereignty addresses who manages and accesses an environment, including personnel, vendors, and support systems. It extends beyond where data is stored.

Q: Why is operational sovereignty commonly overlooked?

A: Many organizations focus primarily on data location and encryption. Access pathways, support personnel, and telemetry flows are often not fully audited.

Q: How do vendors impact sovereignty posture?

A: Vendors and managed service providers may have access to sensitive systems or metadata. Their legal jurisdictions and operational practices can affect overall sovereignty compliance.

Q: Why are telemetry and metadata important?

A: Even if primary data remains local, telemetry and metadata may cross jurisdictional boundaries. These flows can create compliance risks if left unmanaged.

Q: What does a strong operational sovereignty model include?

A: It includes mapped access pathways, documented jurisdictional controls, audited vendor access, and visibility into all telemetry and metadata flows.

Alex Zinin is VP/GM, Managed Service Providers, at Commvault.

More related posts


Thumbnail-Digital-Sovereignty-4

Sovereign Data You Can’t Recover Isn’t Actually Sovereign

Read more about Sovereign Data You Can’t Recover Isn’t Actually Sovereign
Thumbnail-Digital-Sovereignty-2

Minimum Viable Sovereignty: Why the Right Posture Isn’t the Same for Every Organization

Read more about Minimum Viable Sovereignty: Why the Right Posture Isn’t the Same for Every Organization
Thumbnail-Digital-Sovereignty-1

You Don’t Have a Sovereignty Strategy. You Have a Residency Policy.

Read more about You Don’t Have a Sovereignty Strategy. You Have a Residency Policy.

Die wichtigsten Erkenntnisse

  • Bei souveränen Architekturen wird häufig der Schwerpunkt eher auf Audits und Zugriffskontrollen als auf die Wiederherstellungsbereitschaft gelegt.
  • Personal für die Wiederherstellung, Backup-Systeme und Modelle zur Verwahrung von Schlüsseln können bei Vorfällen zu Lücken in der Souveränität führen.
  • Einheitliche Kontrollmaßnahmen in Primär- und Wiederherstellungsumgebungen sind unerlässlich.
  • Eine auf staatliche Anforderungen ausgerichtete Widerstandsfähigkeit erfordert erprobte Wiederherstellungsverfahren unter realistischen Bedingungen.

Picture the moment. The attack has already happened. The incident response team is assembling. Someone must decide which systems come back first, in what order, using the correct recovery points.

And then someone realizes: The personnel with recovery system access are based in a different country. Worse, the recovery environment itself (hosted in a cloud region, a partner datacenter, or a secondary site) was never subject to the same sovereignty controls as the primary data.

The practice wasn’t subject to the same sovereignty controls as the primary data. The regulator is asking for status. The clock is running.

This is the scenario most sovereign architectures were not designed for – and the one the „„Digital Sovereignty Readiness Report““ calls out directly: most sovereign applications are designed for the audit, not the incident.

Most sovereign applications are designed for the audit, not the incident. The difference becomes visible at the worst possible moment.

Der blinde Fleck der Erholung in der Architektur der Staatsfinanzen

Sovereignty programs are built around access control – who can reach the data, under what authority, through what pathway. That architecture is necessary. It is not sufficient. And it connects directly to the operational sovereignty gaps die im dritten Beitrag dieser Reihe beleuchtet wurden: If the people who run your environment operate outside your sovereignty boundary, that problem doesn’t disappear during an incident. It becomes the problem.

What access control leaves unanswered is the harder question: What happens after an incident, when recovery is not just a technical operation but a legally constrained one?

A ransomware attack on a regulated European organization doesn’t simply create a recovery problem. It creates a recovery problem that must be solved within a jurisdiction, using personnel with appropriate authorizations, against recovery points that can be demonstrated to be clean and uncompromised.

The sovereign architecture designed to protect the data can make recovery harder if resilience wasn’t built into the original design.

Die spezifischen Ausfallarten

The ways sovereign recovery architectures fail are predictable – and common:

  • Recovery personnel outside the sovereignty boundary. The engineers who know the recovery systems may operate in a different jurisdiction. Under pressure, using them is the path of least resistance. It is also a sovereignty violation at the moment it is least convenient to have one.
  • Backup infrastructure without matching controls. Primary sovereign environments are carefully controlled. Backup infrastructure – particularly older or secondary environments – is frequently not subject to the same sovereignty requirements. If recovery points are stored or processed outside the boundary, compliant recovery is not available from compliant infrastructure.
  • Key custody under crisis conditions. Hold-your-own-key arrangements are designed for normal operations. Under crisis conditions – with primary systems compromised and time pressure acute – the key custody model that works in a routine maintenance window may become an obstacle to recovery. If this hasn’t been tested, it’s an assumption, not a control.
  • Cross-environment governance gaps. Organizations operating across multiple sovereign tiers – which is most of them – often have strong controls in primary environments and weaker controls in secondary environments that are also part of the recovery path. Consistency across the full estate is what auditors will look for. Gaps in secondary environments become visible exactly when consistency matters most.

Warum staatliche Kontrollen die Erholung erschweren können

The same controls that make a sovereign environment defensible to an auditor can make it harder to recover from. Data movement restrictions that prevent unauthorized exfiltration also constrain recovery orchestration. Key custody arrangements that ensure no provider can access your data without authorization also add friction when you need to restore quickly.

None of this means these controls are wrong. It means they have to be designed with recovery in mind from the start – not added to an architecture where recovery was an afterthought. This is the core of the minimum viable sovereignty principle: Calibrating controls to actual requirements includes recovery requirements, not just access control requirements.

Was „Sovereignty-Ready Resilience“ erfordert

  • Clean recovery validation. Proving that recovery points are free from compromise before restoring to production – not just recent, but uncompromised. In a ransomware scenario, a recent backup may itself be compromised. The ability to identify and restore from a known-clean recovery point, validated before it’s needed, is a sovereignty requirement, not just a disaster recovery requirement.
  • Cross-environment governance. Consistent sovereignty controls and audit evidence across the full estate – not just the primary sovereign deployment. Every environment in the recovery path must meet the same requirements as the primary environment.
  • Tested under realistic conditions. Regular exercises that validate recovery under the conditions that will actually exist during an incident: the legal constraints that apply, the personnel who are available, the recovery points that are clean. An annual disaster recovery test that doesn’t account for sovereignty constraints is not a sovereignty-ready exercise.

Die Frage, die Sie Ihrer „Sovereignty Review“ hinzufügen sollten

There is a direct way to assess whether your recovery architecture meets the same sovereignty requirements as your primary data environment: Ask it as a question and require an honest answer.

Can you recover your sovereign data, cleanly, within defined tolerances, using personnel operating within your sovereignty boundary, right now – under real conditions, not a controlled exercise?

For most organizations, the honest answer reveals a gap. The organizations that find it now – before the incident – will be best prepared with evidence when the regulator asks for it. The ones that don’t will be building it under pressure, in front of the people they least want to disappoint.

The „„Digital Sovereignty Readiness Report“““ enthält eine Frage zur Bewertung der Architektur für die direkte Wiederherstellung.

FAQs

Q: Why is recovery important to digital sovereignty?

A: Sovereignty is incomplete if organizations cannot recover data within the same legal and operational boundaries used to protect it.

Q: What are common sovereign recovery failures?

A: Common failures include recovery personnel operating outside the sovereignty boundary, backup infrastructure lacking matching controls, and inconsistent governance across environments.

Q: How can key custody complicate recovery?

A: Hold-your-own-key models strengthen security during normal operations, but they can slow recovery efforts during incidents if not properly tested.

Q: What is clean recovery validation?

A: Clean recovery validation confirms that recovery points are free from compromise before systems are restored. This is especially important in ransomware scenarios.

Q: How should organizations test sovereignty-ready resilience?

A: They should conduct realistic exercises that account for legal constraints, operational availability, and validated recovery points – not just standard disaster recovery testing.

Alex Zinin is VP/GM, Managed Service Providers, at Commvault.

More related posts


Thumbnail-Digital-Sovereignty-2

Minimum Viable Sovereignty: Why the Right Posture Isn’t the Same for Every Organization

Read more about Minimum Viable Sovereignty: Why the Right Posture Isn’t the Same for Every Organization
Thumbnail-Digital-Sovereignty-3

The Pillar Most Sovereignty Strategies Forget

Read more about The Pillar Most Sovereignty Strategies Forget
Thumbnail-Digital-Sovereignty-1

You Don’t Have a Sovereignty Strategy. You Have a Residency Policy.

Read more about You Don’t Have a Sovereignty Strategy. You Have a Residency Policy.

Die wichtigsten Erkenntnisse

  • Bei der „Minimum Viable Sovereignty“ (MVS) geht es darum, für die jeweiligen Workloads das richtige Maß an Kontrolle anzuwenden.
  • Die Gleichbehandlung aller Workloads kann zu unnötiger Komplexität und Kosten oder zu unzureichendem Schutz führen.
  • Unternehmen lassen sich in der Regel in drei Souveränitätsprofile einteilen: „True Sovereign“, „Regulated Enterprise“ und „Hybrid Multi-Cloud“.
  • Eine einheitliche Governance über gemischte Umgebungen hinweg ist eine der größten betrieblichen Herausforderungen.

There is a version of the digital sovereignty conversation that leads organizations somewhere expensive, operationally burdensome, and – if they’re being honest – further than their actual obligations require. Maximum sovereignty sounds responsible. In practice, it’s often a miscalibration.

There is an equally common version that leads somewhere dangerously thin – controls that satisfy a checklist but wouldn’t survive an audit, an incident, or a regulator who has stopped accepting documented intent as proof of demonstrated control.

The organizations that get sovereignty right tend to do something more rigorous and more practical than either extreme: They ask what they actually owe, to whom, and for what. Then they build to that standard – no more, no less.

This is the discipline of MVS, introduced in the „„Digital Sovereignty Readiness Report““  and developed in full here.

MVS isn’t a shortcut. It’s a recognition that the goal is the right level of control, applied consistently, across every workload that requires it.

Nicht alle Workloads sind gleich

The starting point for an MVS approach is workload classification – and most organizations skip it entirely.

A trading system processing regulated financial data carries fundamentally different sovereignty obligations than an internal HR collaboration tool. A database holding personal data of EU citizens is subject to a different legal and regulatory regime than a development environment running anonymized test data.

Treating all of these identically – either by applying maximum sovereign controls across the board or by assuming a single deployment model covers everything – is how organizations end up either over-engineered or under-protected.

The right question before any deployment decision: What does this workload require across each of the four sovereignty pillars? The Readiness Report includes a self-assessment structured around exactly that question.

The Three Profiles – and What They Actually Need

Regulierte Unternehmen lassen sich in drei erkennbare Profile einteilen, die jeweils unterschiedliche Hauptantriebsfaktoren und Investitionsprioritäten aufweisen.

  • The True Sovereign. Government agencies, defense contractors, and critical national infrastructure operators. For these organizations, sovereignty is not a compliance requirement – it is an operational mandate. Maximum control over every dimension of the technology stack is often legally required, and the cost tradeoffs are accepted because the alternative is not.
  • The Regulated Organization. Financial services firms, healthcare organizations, energy companies. These organizations face binding requirements from DORA, NIS2, GDPR, and sector-specific frameworks. Compliance obligations may also map to EU certification schemes – including EUCS, EUCC, BSI C5, and SecNumCloud – depending on sector and deployment context.

on-negotiable in certain areas – particularly around data residency, operational access controls, and recovery within jurisdictional boundaries. But not every workload carries the same obligation.

  • The Hybrid Multi-Cloud Organization. Organizations with existing hyperscaler investments facing increasing sovereignty pressure from customers, regulators, or procurement requirements. Their challenge is not wholesale migration – it’s layering sovereign controls onto a mixed estate and maintaining consistent governance across it.

Die Kosten einer falschen Kalibrierung

Over-engineering sovereignty creates its own operational risks. Organizations that apply maximum sovereign controls to workloads that don’t require them absorb cost and complexity that serves no regulatory or business purpose.

Under-engineering is the more common failure mode, and the more dangerous one. It typically doesn’t show up until the audit arrives – or, more seriously, until an incident occurs and recovery becomes a legally constrained problem. (That failure mode is the subject of the vierten Beitrags dieser Reihe) umfasst.

Ein praktischer Ausgangspunkt

Ein MVS-Ansatz umfasst drei Schritte:

  1. Classify workloads by their actual sovereignty requirements across each pillar – don’t start with deployment models.
  2. Ordnen Sie jede Workload-Klasse der Bereitstellungsstufe zu, die diese Anforderungen erfüllt, und zwar über das gesamte Spektrum hinweg – von Regionen öffentlicher Hyperscaler über souveräne Public Clouds bis hin zu lokal verwalteten Umgebungen.
  3. Govern the resulting mixed estate consistently – controls, audit evidence, and recovery capabilities must be demonstrable across the full environment, not just the most-sovereign tier.

The third step is where most programs struggle. Maintaining consistent sovereignty controls across a mixed estate is an operational governance challenge – and specifically the domain of Operational Sovereignty – dem Thema des dritten Beitrags dieser Reihe, der Säule, die in den meisten Strategien erst nachträglich berücksichtigt wird.

Nutzen Sie die Selbsteinschätzung im„„Digital Sovereignty Readiness Report“““, um Ihre aktuelle Situation in allen vier Säulen zu ermitteln.

FAQs

Q: What is minimum viable sovereignty (MVS)?

A: MVS is the practice of applying sovereignty controls based on actual business and regulatory needs. It is intended to help avoid both over-engineering and under-protection.

Q: Why is workload classification important?

A: Different workloads carry different regulatory and operational obligations. Classifying workloads helps organizations apply the appropriate level of sovereignty controls.

Q: What are the three common sovereignty profiles?

A: The three profiles are true sovereign organizations, regulated organizations, and hybrid multi-cloud organizations. Each has distinct operational and compliance requirements.

Q: What risks come from over-engineering sovereignty?

A: Excessive controls can increase operational complexity and costs without delivering meaningful compliance or business value.

Q: Why do mixed environments create governance challenges?

A: Organizations often operate across multiple cloud and infrastructure models. Maintaining consistent controls, audit evidence, and recovery standards across all environments is difficult.

Ruben Renders is Solutions Director, MSP, at Commvault.

More related posts


Thumbnail-Digital-Sovereignty-4

Sovereign Data You Can’t Recover Isn’t Actually Sovereign

Read more about Sovereign Data You Can’t Recover Isn’t Actually Sovereign
Thumbnail-Digital-Sovereignty-3

The Pillar Most Sovereignty Strategies Forget

Read more about The Pillar Most Sovereignty Strategies Forget
Thumbnail-Digital-Sovereignty-1

You Don’t Have a Sovereignty Strategy. You Have a Residency Policy.

Read more about You Don’t Have a Sovereignty Strategy. You Have a Residency Policy.

Die wichtigsten Erkenntnisse

  • Data residency addresses where data is stored, but digital sovereignty also requires control over access, operations, and proper understanding of jurisdictional implications.
  • Operational sovereignty is often the weakest and least-audited part of most sovereignty programs.
  • A complete sovereignty posture depends on four pillars: data locality, technological sovereignty, operational sovereignty, and jurisdictional sovereignty.
  • Sovereignty is not binary; organizations must define a posture aligned to their regulatory and operational obligations.

Here is a question worth sitting with: When your organization made its sovereignty decision, what exactly did it decide?

For most, the answer is some version of the same thing. Pick a region. Move the workloads. Choose a cloud provider with data centers in-country. Check the box. The question of where data lives was answered, and the sovereignty conversation was considered closed.

But it wasn’t closed. It had barely started.

Data residency answers one question: Where? Digital sovereignty asks three more – who, how, and under what conditions?

The conflation of residency with sovereignty is understandable. Hyperscalers have made region selection feel like a sovereignty decision. Compliance checklists ask where data is stored. Regulatory guidance, at least in its earlier iterations, focused heavily on geography.

Choosing a sovereign cloud region is a real thing – it matters, it has operational implications, and it’s a necessary first step. But it is only a first step. And most organizations stopped there.

What Residency Doesn’t Answer

Think of it this way: Choosing a sovereign cloud region is like buying a safe. It tells you where your valuables are stored. It says nothing about who has a copy of the combination, who manufactured the safe, which country’s laws govern the manufacturer, or whether you can open it if compelled to.

Region selection answers one question. Three more remain entirely open – and these are the questions regulators, procurement committees, and auditors are now asking with increasing precision:

  • Who can operate your environment, and from where? Whether your cloud provider’s support personnel are subject to foreign jurisdiction is a sovereignty question that data residency cannot resolve. A routine maintenance window performed by a support engineer in a different legal jurisdiction is an access pathway your residency policy doesn’t cover. This is the domain of Operational Sovereignty – the hardest pillar to audit and the most commonly overlooked.
  • Under what legal regime can your data be accessed? A foreign technology provider operating infrastructure in-country does not automatically remove the reach of their home jurisdiction’s law. The extraterritorial reach of foreign legal regimes is a risk that geography alone cannot eliminate.
  • Can you recover your data if something goes wrong? Most sovereignty programs are built around access control. Very few address recovery – whether your data can be restored cleanly, within defined tolerances, by personnel who operate within your sovereignty boundary. That gap is where sovereignty postures most commonly fail under real conditions.

The Framework that Fills the Gap

A complete sovereignty posture spans four interdependent pillars. The „„Digital Sovereignty Readiness Report““ – available at readiverse.com – walks through each in full. In brief:

  • Data locality addresses where data and metadata actually travel.
  • Technological sovereignty covers control over encryption, key custody, and architecture portability.
  • Operational sovereignty covers who runs the environment and from where.
  • Jurisdictional sovereignty establishes the legal framework governing and affecting all of the above.

No single pillar is sufficient. A strong data locality posture with weak operational controls is not sovereignty – it is residency with unexamined risk.

What makes the framework useful is not its complexity. It’s the questions it generates. When an organization maps its current posture against all four pillars for the first time, it almost always finds gaps it didn’t know were there – not because the controls are absent, but because the questions were never asked.

Sovereignty Is a Sliding Scale

One more thing worth naming: Sovereignty is not a binary state. There is no certification that grants it and no single deployment model that guarantees it. It is a posture – a set of deliberate, auditable decisions. And the right level of that posture varies by organization, by workload, and by what you actually owe regulators and customers.

That calibration is what minimum viable sovereignty is about – the subject of the second post in this series.

Regulatory confidence is built long before the audit itself – through clearly defined requirements, not assumptions tied to geography.

Download the „„Digital Sovereignty Readiness Report““ for the four-pillar framework and a practical self-assessment tool.

FAQs

Q: What is the difference between data residency and digital sovereignty?

A: Data residency focuses on where data is physically stored. Digital sovereignty goes further by addressing who can access the data, how systems are operated, and exposure to which jurisdictions may create legal risk.

Q: Why is region selection not enough for sovereignty?

A: Choosing a cloud region only addresses geography. It does not resolve issues related to operational access, legal risks exposure, or recovery capabilities.

Q: What are the four pillars of digital sovereignty?

A: The four pillars are data locality, technological sovereignty, operational sovereignty, and jurisdictional sovereignty. Together, they create, what we believe, is a more complete framework for assessing sovereign readiness.

Q: Why is operational sovereignty difficult to manage?

A: Operational sovereignty involves monitoring who can access systems, where they operate from, and under which legal regime. These controls are harder to audit than simple data location requirements.

Q: Is digital sovereignty a fixed certification?

A: No. Sovereignty is an ongoing posture based on deliberate, auditable decisions that vary by organization, workload, and regulatory environment.

Ruben Renders is Solutions Director, MSP, at Commvault.

More related posts


Thumbnail-Digital-Sovereignty-3

The Pillar Most Sovereignty Strategies Forget

Read more about The Pillar Most Sovereignty Strategies Forget
Thumbnail-Digital-Sovereignty-4

Sovereign Data You Can’t Recover Isn’t Actually Sovereign

Read more about Sovereign Data You Can’t Recover Isn’t Actually Sovereign
Thumbnail-Digital-Sovereignty-2

Minimum Viable Sovereignty: Why the Right Posture Isn’t the Same for Every Organization

Read more about Minimum Viable Sovereignty: Why the Right Posture Isn’t the Same for Every Organization