AlgoMaster Logo

Availability

High Priority11 min readUpdated September 16, 2026
AI Mock Interview

Practice this topic in a realistic system design interview

Listen to this chapter
Unlock Audio

Premium Video

This video is available to premium subscribers only

Unlock Full Access

A system can handle plenty of traffic and still fail its users if one broken component takes the whole thing down. Handling more load and staying online during failures are related problems, but they are not the same problem.

Availability is about the second one: making sure the service is there when users need it. This chapter covers what availability means, how to measure it, and the main approaches to keep a system available even when something fails. We will use an online store as the example throughout.

1. The Problem

Suppose our online store runs on one application server connected to one database. Customers can browse products, add items to their carts, and place orders.

Everything works until the application server crashes. The database is still healthy, but customers can no longer use the store.

We need a way to keep serving customers when part of the system stops working. That is where availability comes in.

2. What Availability Means

Availability measures how often users can access and use a system.

Scalability is about handling more load, while availability is about making sure the service is there when users need it. A store that can process ten thousand orders per minute has solved the first problem. If a single server crash takes it offline for an hour, it has not solved the second.

Availability is closely related to reliability, but they are not the same thing. Reliability measures whether the system continues to work correctly over time. For example, an online store might still be available and accepting orders, but if a bug calculates the wrong order total, the system is available but not reliable.

Scroll
PropertyQuestion It AnswersStore Example
ScalabilityCan the system handle more load?Handling ten times more orders during a sale
AvailabilityCan users reach and use the system?Customers can still check out when a server crashes
ReliabilityDoes the system behave correctly over time?Every order total is calculated correctly

3. Measuring Availability

The basic formula to measure availability is simple: uptime divided by total time, where total time is uptime plus downtime.

Availability = Uptime / (Uptime + Downtime)

For example, suppose our store is available for 364 days and down for 1 day over the course of a year. We divide 364 by 365 and multiply by 100, which gives us roughly 99.73% availability.

Availability = 364 / 365 = 99.73%

But that number alone does not tell the full story. We also need to define what part of the system we are measuring and over what period of time.

For example, users might still be able to browse products even if checkout is completely unavailable. A single availability number for the entire store can hide that failure.

Scroll
OperationAvailability This MonthWhat Users Experienced
Browse products99.99%A few minutes of errors
Add to cart99.95%About 20 minutes of errors
Checkout99.5%Over 3 hours of failed orders
Whole store (averaged)99.8%Hides the checkout outage

Measuring per operation, over a stated window, shows where the real problem is. Checkout is the operation the business cares most about, and it is the one with the worst number.

4. The Nines of Availability

You will often hear availability described in terms of "nines." Three nines means 99.9% availability, which allows roughly 8 hours and 46 minutes of downtime per year. Four nines, or 99.99%, brings that down to about 53 minutes. Five nines allows just over 5 minutes, while six nines allows only around 32 seconds.

Loading simulation...

Scroll
AvailabilityDowntime per YearDowntime per MonthDowntime per Week
99% (two nines)3.65 days7.3 hours1.68 hours
99.9% (three nines)8.76 hours43.8 minutes10.1 minutes
99.99% (four nines)52.6 minutes4.38 minutes1.01 minutes
99.999% (five nines)5.26 minutes26.3 seconds6.05 seconds
99.9999% (six nines)31.5 seconds2.63 seconds0.6 seconds

Each additional nine reduces the allowed downtime by roughly a factor of ten. That gives us much less time to detect failures and recover from them. At four nines, a single incident that takes an hour to fix has already used the whole year's budget.

So higher availability targets come at a cost. They usually require more redundancy, spare capacity, automation, monitoring, testing, and faster recovery mechanisms.

TargetTypical Requirements
99.9%A second server, basic monitoring, a person who can respond within a few hours
99.99%Automatic failover, health checks, on-call rotation, tested recovery procedures
99.999%Redundancy at every layer, multiple locations, automated recovery with no human in the loop

5. Components in Series

Now let's look at the path a request takes through the system. Suppose every request has to pass through a web server, an application server, and a database.

For the request to succeed, all three components need to be available. We call this a series dependency.

If the failures are independent and each component has 99.9% availability, we multiply 0.999 by itself three times. That gives us an overall availability of roughly 99.7% for the complete request path.

Overall = 99.9% × 99.9% × 99.9% = 99.7%

The effect compounds with every component added. Five components at 99.9% each give a combined figure of about 99.5%, which is almost 44 hours of downtime per year, compared with under 9 hours for a single component.

The important idea is that every required component introduces another point of failure. So even when each individual component is highly available, the availability of the entire system can be lower.

6. Components in Parallel

Now suppose we have two application servers, and either one can handle the request. In this case, the servers are working in parallel.

If each server has 99.9% availability, then each is down 0.1% of the time. Assuming their failures are independent, the probability that both servers are down at the same time is 0.1% multiplied by 0.1%, which is 0.0001%.

Failure probability = 0.1% × 0.1% = 0.0001%

Availability = 100% - 0.0001% = 99.9999%

That gives the pair a theoretical availability of 99.9999%, or six nines.

But there are a couple of important assumptions here. Traffic must be able to reach the healthy server, and that server must have enough capacity to handle the extra load. And this availability applies only to these two servers, not to the database or the rest of the request path.

Series and parallel are mirror images. In series, the system fails when any one component fails, so combined availability is lower than the weakest part. In parallel, the system fails only when every redundant component fails together, so combined availability is higher than any single part.

7. Why Independence Matters

That independence assumption is important. Two servers can still fail at the same time if they depend on the same power source, receive the same faulty deployment, or share a network dependency.

The failures that hit both servers at once come from a few recurring sources.

Scroll
Failure CategoryExamplesHits Both Servers When
HardwareA drive, server, switch, or power supply stops workingThey share a rack, power feed, or switch
SoftwareA bug, memory leak, deadlock, or cascading failureThey run the same faulty release
NetworkPacket loss, latency spikes, a partition, or DNS failingThey sit behind the same network path or DNS entry
Human errorA wrong config value, a bad deployment, an accidental deletionThe same change is applied to both

Hardware failures deserve a specific note. At small scale, a server failure feels rare. At large scale, it is routine. A fleet of thousands of servers sees drives and power supplies fail every week, so the design has to assume that any single machine can disappear at any time.

So simply adding more machines is not enough. That is the core idea behind redundancy: provide another way to handle the requests, and design the system so that it can survive the failure of one server or node completely.

8. Active-Passive Redundancy

One common way to provide redundancy is an active-passive setup. One component handles traffic, while another stays ready as a standby.

For example, in our store's database, the primary database accepts writes and replicates its changes to the standby.

If the primary fails, a monitoring system detects the failure, promotes the standby to become the new primary, and redirects application connections to it. This process is called failover.

Keeping a single active writer makes data changes easier to coordinate. There is one place that accepts writes, so there is no question about which copy holds the latest order.

The trade-off is that failover takes time. Detecting the failure, promoting the standby, and reconnecting the application can cause a brief interruption for users.

Pros

  • One clear writer: Data changes are easy to coordinate because only the primary accepts writes.
  • Simple to reason about: There is one active copy and one backup.
  • Lower resource use: The standby does not need to serve production traffic.

Cons

  • Failover takes time: Detection, promotion, and reconnection add up to an interruption.
  • Standby may be untested: A backup that never serves real traffic can surprise you during an outage.
  • Split-brain risk: If the monitor promotes the standby while the old primary is still alive, both may accept writes.

Standby Readiness

How quickly we can recover also depends on how ready the standby is.

Scroll
Standby TypeStateFailover TimeCost
Cold StandbyNot running, may need data restoredMinutesLowest
Warm StandbyRunning and configured, not serving trafficSeconds to minutesMedium
Hot StandbyRunning, data in sync, ready to serveSecondsHighest

A cold standby is not running, so during a failure we first need to start it and may also need to restore or load the latest data.

A warm standby is already running and configured, but it still needs some preparation before it can take over, such as catching up on recent changes and being promoted.

A hot standby stays up to date and is ready to serve almost immediately, which makes failover much faster.

The trade-off is cost. The more prepared the standby is, the more infrastructure and resources we need to keep running.

And just having a standby is not enough. We also need to test it under realistic production load. A backup database is not very useful if, during an outage, we discover that it cannot handle the traffic.

9. Active-Active Redundancy

For the application servers, we can use an active-active setup instead. Both servers handle requests through a load balancer, so if one fails, the other is already running and can continue serving traffic immediately.

But the surviving server still needs enough spare capacity to handle the extra load. If both servers are already close to their limits, losing one could overload the other.

We also need to make sure either server can handle any request. That means shared state, such as cart and session data, should live outside the application servers in a shared data store.

This is why stateless application servers work well with active-active setups. Any healthy server can pick up the next request, because nothing the request needs is stuck on the server that failed.

Pros

  • No failover delay: The surviving server is already serving traffic.
  • Every node is tested under real load: There is no untested backup.
  • Better use of hardware: No machine sits idle.

Cons

  • Needs spare capacity: Each server must be able to absorb the others' share.
  • Needs shared state: Session and cart data must live outside the servers.
  • Harder for databases: Multiple active writers must coordinate changes.

Health Checks and the Load Balancer

The load balancer uses health checks to decide which servers should receive traffic. If a server becomes unhealthy, the load balancer stops sending new requests to it.

But detection is not instant. Health checks run every few seconds, and a server is usually marked unhealthy only after several checks fail in a row. Some requests already in progress, or sent during that window, may still fail.

A good health check should verify that the server can actually handle requests, not just that its process is still running.

And finally, the load balancer itself also needs redundancy. Otherwise, it becomes a single point of failure that can make all of the healthy application servers unreachable.

Managed cloud load balancers are built to run redundantly, so this is handled for you on AWS, Google Cloud, and Azure. In a self-managed data center, teams often run two load balancer instances that share a virtual IP address, so the standby can take over the address when the active one fails.

10. Spreading Across Locations

So far, all of our redundant components could still be running in the same building. That means a facility-wide power or network outage could take them all down at once.

To reduce that risk, we can spread the system across multiple availability zones within the same cloud region. For even greater separation, we can deploy across multiple regions.

Scroll
LevelWhat It IsProtects AgainstLatency Between Copies
Availability ZonesSeparate data centers in the same region, connected by fast linksA single data center losing power or networkLow, a few milliseconds
RegionsSeparate geographic areas, such as US East and EuropeRegional disasters and region-wide outagesHigher, tens of milliseconds or more
Multiple cloud providersDifferent providers, such as AWS and Google CloudA provider-wide outageVariable

But wider separation introduces new trade-offs. Replicating data to a distant region adds network latency, especially if we wait for the remote replica before confirming a write. Asynchronous replication avoids that delay, but a regional failure could leave the backup missing some recent changes.

Scroll
ReplicationHow It WorksData Loss on FailoverEffect on Writes
SynchronousThe write succeeds only after the remote replica confirms itNoneEvery write waits for the round trip to the other region
AsynchronousThe write is confirmed immediately and the replica catches up laterPossible, the last few seconds or minutesNo added latency

Using multiple cloud providers adds another layer of isolation, but it also increases operational complexity. Each provider has its own tooling, networking, and failure behavior, and the team has to be fluent in both.

The right level of redundancy depends on the failures we need to survive and how much cost and complexity we are willing to take on. For most applications, running across two or more availability zones is the sensible default. Multi-region is worth it when a regional outage would be unacceptable or when users are spread across the globe.

11. Handling Overload with Queues

Failures can also happen because the system is overloaded, not because a component broke.

Suppose a big sale causes a sudden spike in orders, and sending confirmation emails starts slowing down checkout. Each checkout request waits for the email provider to respond, and when the provider slows down under the spike, so does every order.

Instead of sending the email during the request, we can save the order first and place an email job in a durable queue. Background workers process those jobs separately, so customers do not have to wait for the email to be sent before seeing their order confirmation.

If a worker fails, unfinished jobs can be retried by another worker. The queue holds the job until a worker confirms it is done, so a crash in the middle of sending does not lose the email.

12. Containing Cascading Failures

Now suppose the product recommendation service of our online store becomes slow. Application requests start waiting on it. As more requests pile up, the application servers run out of threads and connections, and the slowdown spreads to other parts of the store, including checkout. This is called a cascading failure.

To limit the impact, we can use timeouts so requests do not wait indefinitely, and a circuit breaker to temporarily stop calling a dependency that is repeatedly failing.

A circuit breaker tracks the calls to a dependency. While calls succeed, it stays closed and lets requests through. Once failures pass a threshold, it opens and rejects calls immediately, without waiting on the slow service. After a cooldown, it lets a few test requests through to see whether the dependency has recovered.

Scroll
ProtectionWhat It DoesStore Example
TimeoutGives up on a call after a fixed waitStop waiting for recommendations after 200 ms
Circuit breakerStops calling a dependency that keeps failing, then probes for recoverySkip the recommendation call entirely while it is down
FallbackReturns a reduced but usable responseShow the product page without recommendations

For recommendations, we can add a fallback that shows the product page without recommendations. Customers can still browse and place orders while the recommendation service recovers.

The recommendation service is still down in this picture. What changed is that its failure stays contained. Checkout, which never needed recommendations, keeps working at full speed.

Summary

Design for components to fail while keeping a usable path available for customers. That is the idea behind every technique in this chapter.

Scroll
TechniqueFailure It HandlesApplied in the Store
Measure per operationA single number hiding a real outageTrack checkout availability separately from browsing
Active-passive with failoverThe primary database going downStandby promoted to primary
Active-active behind a load balancerAn application server going downThe other server keeps serving
Health checksTraffic sent to a broken serverUnhealthy servers removed from rotation
Multiple zones or regionsA whole facility or region going downStore runs in more than one location
Durable queuesOverload from slow background workConfirmation emails sent by workers
Timeouts, circuit breakers, fallbacksA slow dependency dragging everything downProduct pages without recommendations

Measure availability around the operations users actually care about. Use redundancy and safe failover for critical components, queues for work that can happen in the background, and timeouts and circuit breakers to prevent failing dependencies from bringing down the rest of the system.

That is the core of availability.

As an exercise, take a simple application and trace what happens when each component fails. Identify which failures can take down the entire service, and think about how you would recover from them.

Quiz

Availability Quiz

10 quizzes