High Availability (HA) and Disaster Recovery (DR)

High Availability (HA) and Disaster Recovery (DR)

Definition: High Availability (HA) is the practice of architecting a system so it keeps operating through localized failures — a dead disk, a crashed process, a whole server or Availability Zone going down — without a noticeable interruption to users. Disaster Recovery (DR) is the broader set of processes, tooling, and policies for restoring service after a catastrophic event that HA alone can’t absorb, such as an entire region going offline, a ransomware attack, or a large-scale data-corruption incident. The two are commonly discussed together because they form a continuum of failure severity rather than separate disciplines: HA is what keeps the lights on for common, expected failures, and DR is the plan for the rarer, larger failures HA wasn’t designed to survive. Both are measured in the same two numbers — RTO and RPO — which is what makes them possible to design, budget, and test against rather than treat as vague aspirations.

How It Works

  • HA within a region: achieved by deploying redundant resources — load balancers, Auto-Scaling groups, database replicas — across multiple isolated Availability Zones (AZs), each with independent power, cooling, and networking, so a single AZ failure doesn’t take the whole system down
  • DR across regions: relies on backups and geographic replication to an entirely different cloud region (e.g., US-East to US-West), so a regional-scale event — a natural disaster, a provider-wide outage, a botched deployment that only a full region rollback can fix — still leaves a recoverable copy of the system elsewhere
  • Recovery Time Objective (RTO): the maximum acceptable downtime before the system must be restored and serving traffic again, driving how much automation and standby capacity a DR plan needs
  • Recovery Point Objective (RPO): the maximum acceptable data loss, measured in time (e.g., losing at most the last 15 minutes of database writes), driving how frequently data must be backed up or replicated
  • Health checks and failover: load balancers and DNS-based failover (e.g., Route 53 health checks) continuously probe endpoints and automatically redirect traffic away from unhealthy instances or entire failed regions, which is what turns “we have a backup region” into an actual working failover rather than a manual, hours-long scramble
  • Replication mode: synchronous replication (every write acknowledged in two places before returning success) gives a near-zero RPO but adds latency and can stall writes if the secondary is unreachable; asynchronous replication is faster and more tolerant of distance but risks losing the last few seconds or minutes of writes if the primary dies before they replicate

Why It Matters

  • Prevents revenue loss and reputational damage when cloud providers experience the hardware failures, network cuts, or regional outages that are a statistical certainty at large enough scale, not a hypothetical edge case
  • Lets a business quantify risk in concrete, budgetable terms — “we can tolerate 5 minutes of downtime and 1 minute of data loss” is a design constraint engineers can build to, “we should be reliable” isn’t
  • Increasingly a compliance and contractual requirement — SLAs, cyber-insurance policies, and regulations in finance and healthcare often mandate documented and tested RTO/RPO targets, not just good intentions
  • Untested DR plans are a common source of catastrophic surprise; the discipline of HA/DR includes regularly exercising failover (game days, chaos engineering drills), not just having a plan that exists only on paper

Under the Hood: RPO, RTO, and What Failover Actually Does

RPO and RTO answer two different questions about the same incident, and conflating them is the most common design mistake: RPO asks “how much data can we afford to lose,” answered entirely by how data is replicated or backed up before the failure happens, while RTO asks “how long can we afford to be down,” answered by how fast the system can detect the failure and redirect traffic after it happens — a tight RPO doesn’t buy you a tight RTO for free, and vice versa. Achieving a near-zero RPO typically means synchronous or near-synchronous replication to a standby, which is expensive in latency and infrastructure; achieving a near-zero RTO typically means a warm or hot standby that’s already running and just needs traffic redirected, rather than a cold environment that needs to be provisioned from scratch. When failover actually triggers, the sequence is: a health check (load balancer probe, database heartbeat, or external monitoring) crosses a failure threshold after some number of consecutive failed checks (to avoid failing over on a single blip), a failover controller or DNS system updates routing to point at the standby, the standby is promoted from replica to primary (which, for a database, means replaying any remaining replication lag and accepting writes for the first time), and clients reconnect — either automatically via a connection string pointing at a stable endpoint, or after their existing connections time out and retry. The gap between “failure detected” and “standby fully promoted and serving” is the real RTO, and it’s almost always dominated by promotion and DNS propagation delay rather than by detection time.

Comparison: Active-Active vs Active-Passive vs Backup-and-Restore

Active-ActiveActive-PassiveBackup-and-Restore
RTOSeconds — both sides already serving trafficMinutes — standby must be promoted and traffic redirectedHours to days — infrastructure must be provisioned from scratch
RPONear-zero, often sub-secondLow, seconds to minutes depending on replication modeHigh, limited by backup frequency (often hours)
CostHighest — full duplicate capacity running at all timesModerate — standby sized smaller or scaled down until neededLowest — pay mainly for storage of backups
ComplexityHighest — conflict resolution, data consistency across regionsModerate — one-way replication, defined promotion pathLowest — restore scripts and periodic snapshot testing
Best fitGlobal, latency-sensitive, revenue-critical systemsMost production systems with a real but bounded downtime toleranceInternal tools, archives, low-criticality systems

Common Pitfalls

  • Implementing complex multi-region active-active architectures when a simple active-passive setup would easily meet the business’s actual RTO and RPO requirements at a fraction of the cost and operational complexity
  • Writing a DR plan and never testing it — the first real failover attempt is a terrible time to discover a stale runbook, an expired credential, or a script that only ever ran once, manually, years ago
  • Confusing backups with replication: backups protect against corruption and accidental deletion (since a corrupted write doesn’t get replicated backward), replication protects against hardware/AZ failure — a system commonly needs both, not one instead of the other
  • Setting an aggressive RTO/RPO target without budgeting the infrastructure cost that target actually requires, near-zero RPO is not free, it’s synchronous replication and the latency/cost that comes with it
  • Treating HA as if it covers DR-scale events, multi-AZ redundancy inside one region does nothing if that entire region goes down, region-level failures need region-level redundancy
  • Forgetting dependent systems in the failover plan — DNS, secrets managers, third-party APIs, certificate authorities — so the “failed over” primary application comes up but can’t actually authenticate or resolve the services it depends on

Code Example

# AWS Route 53 DNS failover: primary region record with automatic health-check-based failover
PrimaryRecord:
  Type: AWS::Route53::RecordSet
  Properties:
    Name: api.example.com
    Type: A
    SetIdentifier: primary-us-east-1
    Failover: PRIMARY
    HealthCheckId: !Ref PrimaryHealthCheck
    AliasTarget:
      DNSName: !GetAtt PrimaryLoadBalancer.DNSName
      HostedZoneId: !GetAtt PrimaryLoadBalancer.CanonicalHostedZoneID

SecondaryRecord:
  Type: AWS::Route53::RecordSet
  Properties:
    Name: api.example.com
    Type: A
    SetIdentifier: secondary-us-west-2
    Failover: SECONDARY
    AliasTarget:
      DNSName: !GetAtt SecondaryLoadBalancer.DNSName
      HostedZoneId: !GetAtt SecondaryLoadBalancer.CanonicalHostedZoneID

PrimaryHealthCheck:
  Type: AWS::Route53::HealthCheck
  Properties:
    HealthCheckConfig:
      Type: HTTPS
      ResourcePath: /healthz
      FailureThreshold: 3
      RequestInterval: 10

Best Practices

  • Define RTO and RPO explicitly per system, in writing, before choosing an architecture, they’re business decisions with cost tradeoffs, not just engineering defaults
  • Test failover regularly — a scheduled “game day” that actually fails traffic over to the standby is the only way to know the plan works before an incident forces the question
  • Automate failover for anything with a tight RTO, a plan that requires a human to be paged, read a runbook, and manually execute steps at 3 AM has a much longer real-world RTO than the number on paper
  • Separate backup strategy from replication strategy, and validate both, an untested backup is a hope, not a recovery mechanism
  • Include every dependency — DNS, secrets, certificates, third-party integrations — in the DR plan, not just the application and its database

FAQ

Is High Availability the same thing as Disaster Recovery? No — HA keeps a system running through smaller, more frequent failures (a disk, a server, an AZ) usually within a single region, while DR is the plan for larger, rarer failures (a whole region, a catastrophic data event) that HA alone isn’t designed to survive.

What’s a realistic RTO/RPO for a typical production web application? There’s no universal number, but many teams land on something like an RTO of 15-60 minutes and an RPO of 1-15 minutes for standard active-passive setups, tightening toward seconds only for systems where downtime or data loss is extremely costly.

Does having backups mean I already have Disaster Recovery? Not by itself — backups are one ingredient, a real DR plan also needs a tested procedure for provisioning infrastructure, restoring data, redirecting traffic, and validating the restored system actually works, within the RTO the business needs.

History

  • HA concepts predate the cloud by decades, mainframe and early Unix clustering (failover pairs, heartbeat monitoring) in the 1980s-90s established the core patterns — redundancy, health checks, automatic failover — that cloud HA still uses today
  • AWS’s introduction of Availability Zones alongside EC2 in the mid-to-late 2000s made multi-AZ redundancy a standard, relatively cheap architectural default rather than something only large enterprises could afford to build themselves
  • DR historically meant offsite tape backups and a secondary data center that sat mostly idle, cloud infrastructure-as-code and cross-region replication turned DR from a slow, manual, expensive process into something that can be automated and tested on a schedule
  • Chaos engineering, popularized by Netflix’s Chaos Monkey around 2011, pushed the industry toward proactively testing HA/DR assumptions by deliberately injecting failures rather than waiting for a real incident to reveal gaps

Example

Deploying a Kubernetes cluster across three AWS Availability Zones provides High Availability — a single AZ outage barely registers, since the load balancer simply stops routing to the affected nodes. Shipping continuous database replication plus daily snapshot backups to a secondary AWS region, with a tested runbook to promote that region to primary, provides Disaster Recovery — the piece that survives a failure the first architecture was never designed to absorb.

Dig deeper