IBM Z and LinuxONE - IBM Z

IBM Z

The enterprise platform for mission-critical applications brings next-level data privacy, security, and resiliency to your hybrid multicloud.


#Servers
#IBMZ
#Enterpriseserver
 View Only

Always Available Does Not Mean Always Recoverable

By Rebecca Levesque posted 07/09/26 06:46 PM

  

Why Crash Consistency and Always On Are Not Enough

“Business continuity refers to an organization's ability to maintain critical business functions, minimize disruption and resume normal operations.” — IBM, Business Continuity Overview.

Those last three words are important: resume normal operations. They do not describe restarting infrastructure, failing over to another site, or recovering hardware. They describe a business outcome. Yet many organizations continue to equate high availability with recoverability, assuming that because systems can restart, the business can automatically resume operations. In reality, availability and recoverability are fundamentally different capabilities.

This distinction is becoming increasingly important as organizations depend on digital services, operate in increasingly interconnected environments, and face growing risks from cyber incidents, human error, and logical corruption.

Always available does not mean always recoverable.

The Promise of High Availability

For decades, organizations have invested heavily in technologies designed to minimize downtime and keep critical systems available. In the IBM Z world, this includes technologies such as GDPS, Parallel Sysplex, Metro Mirror, Global Mirror, HyperSwap, and continuous availability solutions. These technologies are exceptional at recovering from infrastructure failures and have transformed expectations around uptime and resiliency.

High availability technologies answer important questions. How quickly can I restart? Can I fail over to another site? Can I recover from a hardware or site outage? How do I minimize downtime? These are critically important questions and organizations should continue investing in these capabilities because they are foundational to operational resilience.

However, they do not necessarily answer the questions that become critical during operational failures, cyber incidents, and recovery events. Which applications and business services are affected? Which datasets are required to recover a critical application? What are the dependencies between jobs, applications, and data? Which recovery copy should be used? Which datasets changed and when? What data is missing, duplicated, or unprotected? What is the impact of restoring this application or dataset? Can the business service resume with trusted and consistent data? Can I recover the business within the required timeframe? These are fundamentally different questions. They are not infrastructure questions. They are recoverability questions.

The Recoverability Gap

One of the most important concepts in resilience is understanding the difference between crash consistency and data consistency. A crash consistent environment may successfully restart applications and systems following an outage, but a data consistent environment allows the business to resume operations with trusted and usable information.

This distinction matters because a system can be available, online, and running while still containing corrupted data, incomplete transactions, deleted information, inconsistent applications, or untrusted business records. From an infrastructure perspective, the recovery may appear successful. From a business perspective, however, the recovery may still be incomplete.

Crash consistency is important, but it is not the same as data consistency.

Replication technologies are extremely effective at moving data from one place to another. They are equally effective at replicating mistakes. Accidental deletion, application defects, incorrect updates, ransomware encryption, malicious insider activity, and logical corruption can all be replicated almost immediately. Organizations may therefore have multiple highly available copies of the same problem.

Independent research has observed that too many organizations assume they are more resilient than they actually are, and only about one-third of senior IT decision makers are extremely confident in their disaster recovery and business continuity plans. The confidence gap is real and demonstrates that availability and recoverability are fundamentally different outcomes.

Availability protects systems. Recoverability protects businesses.

Real World Examples

The following examples are not IBM Z or mainframe incidents, and that is precisely why they are relevant. The principles of recoverability are technology independent. Whether the platform is cloud, distributed, or mainframe, organizations face many of the same challenges: accidental deletion, logical corruption, software defects, operational mistakes, ransomware, restoring trusted data, and resuming business services.

GitLab's well-publicized outage demonstrated that having multiple copies of data does not necessarily mean that organizations can recover quickly. The UniSuper and Google Cloud incident highlighted how accidental deletion can still create significant business disruption in highly resilient environments. Knight Capital demonstrated how software defects and incorrect deployment can rapidly become business failures. The CrowdStrike and Delta Air Lines disruptions showed how quickly technology failures can cascade into broad operational impacts and that restoring systems does not automatically restore business operations.

These examples illustrate an important point: availability alone does not guarantee recoverability. The challenge is often not restarting infrastructure, but understanding what was affected, determining what data can be trusted, and restoring the business to a known good state.

Recoverability is a business problem, not a platform problem.

Business Continuity Is About Outcomes

IBM defines business continuity as the ability to maintain critical business functions, minimize disruption, and resume normal operations. Likewise, ISO 22301 defines business continuity as the capability of an organization to continue the delivery of products or services at predefined acceptable levels following a disruptive incident.

Notice what these definitions emphasize: business services, products, outcomes, and recovery. They do not focus solely on infrastructure availability. The business does not recover because storage comes back online. The business recovers when applications can run, data can be trusted, transactions are correct, and critical business services can resume. This is the very definition of data consistency.

Why This Matters Even More in the AI Era

As organizations increasingly adopt AI, automation, and autonomous operations, the need for trusted operational resilience becomes even more important. AI accelerates decision-making, automates operational processes, and can help organizations manage increasingly complex environments. Unfortunately, it can also accelerate mistakes, propagate bad data faster, and increase the impact of operational errors.

IBM Senior Vice President and Chief Commercial Officer Rob Thomas has consistently emphasized that organizations should assume disruption will occur and focus on resilience, trusted recovery capabilities, and business continuity. In an AI-driven world, the question becomes even more important: can we trust the data and operational intelligence on which our business and AI systems depend?

This challenge becomes even more significant as experienced professionals retire and organizations bring new talent onto the platform. Operational knowledge, application dependencies, and recovery procedures increasingly need to be captured, understood, and operationalized rather than residing solely in the minds of a few experts.

Strong operational resilience and recoverability are becoming prerequisites for the successful adoption of AI.

Where High Availability Ends and Recoverability Begins

Business Event

GDPS / High Availability

What the Business Still Needs

Business Objective

Disk Failure

Minimal additional context

Restart systems

Site Failure

Minimal additional context

Resume infrastructure services

Human Error Deletes Data

Replicated

Trusted recovery point

Recover trusted data

Application Corruption

Replicated

Affected applications and dependencies

Restore usable data

Bad Software Deployment

Replicated

Recovery point and impact analysis

Reverse business impact

Ransomware Encryption

Replicated

Trusted copies and recovery sequence

Resume operations

Missing or Duplicate Backups

No visibility

Recovery readiness insight

Establish confidence

Critical Business Service Recovery

Partial

Application, job, and dataset context

Resume the business

The matrix illustrates an important reality. Infrastructure availability and business recoverability are complementary capabilities, but they solve different problems. High availability technologies are designed to keep systems running and recover from infrastructure failures. Business recovery, however, requires additional context and intelligence. Organizations must understand which applications are affected, which datasets are required, how applications and data are related, which recovery points can be trusted, and whether critical business services can resume operation.

The Question That Matters

For years the resilience conversation focused on one question: How quickly can I restart? Today, organizations increasingly need to answer a different question: Can I recover the business and resume operations with trusted data? That is the recoverability gap and why always available does not mean always recoverable. Availability protects systems. Recoverability protects businesses. High availability keeps systems running. Data consistency allows the business to resume. Recoverability requires both.

References

IBM. Business Continuity Overview.

IBM Institute for Business Value. The Cyber Resilient Organization.

IBM Institute for Business Value. Research on AI and operational resilience.

ISO 22301. Business Continuity Management Systems.

Deloitte. Business Continuity and Crisis Management Services.

IT Pro. Too Many Organizations Assume They're More Resilient Than They Actually Are.

GitLab. Postmortem of Database Outage.

Google Cloud and UniSuper incident reports.

Public reporting on the CrowdStrike and Delta Air Lines operational disruptions.

Knight Capital post-incident analyses.

Rob Thomas, IBM. Thought leadership on resilience, AI, and trusted recovery.

0 comments
4 views

Permalink