IBM Z and LinuxONE - IBM LinuxONE Ecosystem

IBM LinuxONE Ecosystem

IBM LinuxONE Ecosystem

Explore IBM LinuxONE ecosystem to partner, learn and connect


#Servers
#IBMLinuxONE
#Enterpriseserver
 View Only

Why IZBR Matters in a World of GDPS and High Availability

By Rebecca Levesque posted 07/10/26 01:09 AM

  

Why Always On Is Only Part of the Recovery Story

For decades, IBM Z organizations have invested in technologies that keep critical systems available. GDPS, Parallel Sysplex, Metro Mirror, Global Mirror, HyperSwap, and other continuous availability technologies have transformed resiliency expectations and significantly reduced downtime resulting from infrastructure failures.

These technologies are essential and remain foundational to operational resilience strategies. They have helped organizations achieve extraordinary levels of availability and have become critical components of business continuity and disaster recovery architectures.

However, as organizations become increasingly dependent on digital services and as cyber threats, human error, and logical corruption continue to grow, another question has emerged:

What happens when the infrastructure is available, but the data can no longer be trusted?

This is where the conversation begins to shift from availability to recoverability.

Always available does not mean always recoverable.

Availability and Recoverability Are Different Problems

High availability technologies answer critically important questions. They help organizations determine how quickly they can restart, whether they can fail over to another site, how they can recover from a hardware or site outage, and how they can minimize downtime. These are essential capabilities and remain foundational to operational resilience.

However, organizations increasingly need answers to a different set of questions:

·       Which business services are affected?

·       What data can be trusted?

·       Which recovery copy should be used?

·       Which backups are missing or duplicated?

·       What changed?

·       Which applications and business processes depend on this data?

·       Can I resume the business with confidence?

·       Can I recover within the required business timeframe?

These are not infrastructure questions. They are recoverability questions.

The distinction is important because availability and recoverability solve different problems. One focuses on restoring systems and infrastructure. The other focuses on restoring trusted business services and enabling the organization to resume normal operations.

Crash Consistency Is Not Data Consistency

One of the most important concepts in operational resilience is understanding the difference between crash consistency and data consistency.

A crash-consistent environment may successfully restart applications and systems following an outage. A data-consistent environment allows the business to resume operations with trusted and usable information.

A system can be available, online, and running while still containing corrupted data, deleted information, incomplete transactions, ransomware-encrypted datasets, or inconsistent applications. From an infrastructure perspective, the recovery may appear successful. From a business perspective, however, the recovery may still be incomplete.

This is why:

Availability protects systems. Recoverability protects businesses.

Replication Can Replicate Problems

Replication technologies are exceptionally effective at moving data from one place to another. They are equally effective at replicating mistakes.

Human error, logical corruption, application defects, malicious activity, and ransomware encryption can all be replicated almost immediately. Organizations can therefore end up with multiple highly available copies of the same problem.

Industry studies have repeatedly shown that many organizations believe they are more resilient than they actually are and that confidence in recovery capabilities often declines significantly when organizations evaluate whether they can truly resume business operations.

The challenge is no longer simply restarting infrastructure. The challenge becomes identifying trusted data, understanding business impact, determining the appropriate recovery point, understanding dependencies, and resuming operations with confidence.

Real World Examples

The following examples are not IBM Z incidents, and that is precisely why they are relevant.

The principles of recoverability are technology independent. Whether the platform is cloud, distributed, or mainframe, organizations face many of the same challenges, including accidental deletion, logical corruption, software defects, ransomware, restoring trusted data, and resuming business services.

GitLab’s widely publicized outage demonstrated that having copies of data does not automatically mean an organization can recover quickly. The UniSuper and Google Cloud incident showed how accidental deletion can create significant business disruption even within highly resilient environments. The CrowdStrike and Delta Air Lines disruptions demonstrated how quickly technology failures can cascade into broad operational impacts and that restoring systems does not automatically restore business operations.

These incidents reinforce an important lesson:

Recoverability is a business problem, not a platform problem.

Business Continuity Is About Outcomes

IBM defines business continuity as the ability to maintain critical business functions, minimize disruption, and resume normal operations. Likewise, ISO 22301 defines business continuity as the capability of an organization to continue the delivery of products or services at predefined acceptable levels following a disruptive incident.

Notice what these definitions emphasize:

·       business services;

·       products;

·       outcomes;

·       recovery.

They do not focus solely on infrastructure availability.

The business does not recover because storage comes back online. The business recovers when applications can run, data can be trusted, transactions are correct, and critical business services can resume.

This is the very definition of data consistency.

Why IZBR Matters

This is where IZBR, the Data Resiliency Manager for Z, complements GDPS and high availability technologies.

GDPS answers an important question:

How do I keep systems available?

IZBR answers a different question:

Can I recover the business?

IZBR provides intelligence that helps organizations understand:

·       what is protected;

·       what is not protected;

·       which backups are missing or duplicated;

·       what has changed;

·       which applications depend on specific data;

·       which business services may be affected;

·       whether recoverability objectives can realistically be achieved.

This relationship is complementary rather than competitive. Availability technologies and recovery intelligence address different aspects of resilience and, together, provide organizations with a more complete understanding of their ability to recover critical business services.

Where High Availability Ends and Recoverability Begins

Business Event

GDPS / High Availability

What the Business Still Needs

Disk failure

Minimal additional context

Site outage

Minimal additional context

Human error deletes data

Replicated

Trusted recovery point

Application corruption

Replicated

Recovery intelligence

Bad software deployment

Replicated

Impact analysis

Ransomware encryption

Replicated

Trusted copies and recovery sequence

Missing or duplicate backups

No visibility

Recovery readiness

Critical business service recovery

Partial

Application and operational context

The matrix illustrates an important reality. Infrastructure availability and business recoverability are complementary capabilities, but they solve different problems.

High availability technologies are designed to keep systems running and recover from infrastructure failures. Business recovery, however, requires additional context and intelligence. Organizations need to understand what data is trusted, what dependencies exist, which recovery points should be used, and whether business services can resume operation.

This is the difference between restoring infrastructure and restoring the business.

The Role of TimeLiner - a component of IZBR

During an operational event or cyber incident, organizations increasingly need to understand what data was used, which processes created it, what applications are affected, how business services are connected, and what downstream impacts may occur.

This knowledge is often difficult to document and frequently resides in the experience of a small number of individuals.

TimeLiner helps make this operational knowledge visible by exposing relationships, dependencies, and operational context that may otherwise remain hidden.

This information has value far beyond recovery events. It can assist organizations with operational resilience, cyber recovery, onboarding New to Z professionals, modernization initiatives, and AI-assisted operations.

The same operational knowledge that helps organizations recover critical business services is often the knowledge required to modernize them successfully.

Why This Matters in the AI Era

As organizations increasingly adopt AI, operational resilience and recoverability become even more important. AI systems depend on trusted data, accurate context, and operational intelligence. Without that foundation, organizations risk accelerating decisions based on incomplete or inaccurate information.

IBM Senior Vice President and Chief Commercial Officer Rob Thomas has repeatedly emphasized in IBM Institute for Business Value research and IBM Think leadership that organizations should assume disruption will occur and focus on resilience, trusted data, and business continuity. He has consistently highlighted that AI is only as valuable as the quality and trustworthiness of the information on which it depends and that resilient organizations are those that can continue operating and adapt in the face of disruption.

This perspective has important implications for operational resilience. AI does not reduce the need for recoverability. In many ways, it increases it. The more organizations depend on automation and AI-driven decision-making, the more important it becomes to understand whether data can be trusted, whether business services can be restored, and whether operational knowledge is available to both people and AI systems.

Modern operational and recovery intelligence therefore become foundational capabilities. They provide the trusted information and context that help people make informed decisions and allow AI systems to augment experience, accelerate understanding, and improve operational effectiveness.

The Question That Matters

For years, the resilience question was:

How quickly can I restart?

Increasingly, organizations are asking a different question:

Can I recover the business and resume operations with trusted data?

GDPS and high availability technologies remain essential because availability matters. However, availability and recoverability solve different, but complementary, problems.

High availability technologies keep systems running and minimize downtime resulting from infrastructure failures. Recoverability capabilities help organizations restore trusted data, understand business impact, and resume critical business services with confidence.

That distinction is becoming increasingly important as organizations adopt AI, modernize applications, and navigate increasingly complex operational environments. Resilient organizations will increasingly be those that invest not only in keeping systems available, but also in understanding whether the business can recover when trusted data is lost, corrupted, or no longer available.

Always available does not mean always recoverable.

Availability protects systems. Recoverability protects businesses.

High availability keeps systems running. Recoverability allows the business to resume. Organizations need both.

References

IBM. Business Continuity Overview.

ISO 22301. Business Continuity Management Systems.

IBM Institute for Business Value. The Cyber Resilient Organization.

IBM Institute for Business Value. Research on AI and operational resilience.

KPMG. From Legacy to Leading: Revolutionizing the Mainframe for the Digital Age.

Deloitte. Business Continuity and Crisis Management Services.

GitLab. Postmortem of Database Outage.

Google Cloud and UniSuper incident reports.

Public reporting on the CrowdStrike and Delta Air Lines operational disruptions.

Rob Thomas, IBM. IBM Think and IBM Institute for Business Value thought leadership on trusted AI, resilience, business continuity, and the need to assume disruption.

0 comments
6 views

Permalink