Availability
The percentage of time a system is operational and accessible. Measured in 'nines', 99.99% availability means about 52 minutes of downtime per year.
What is Availability?
In short
Availability is the percentage of time a system is up and able to serve requests. It is usually written in "nines": 99.9% allows about 8.8 hours of downtime per year, 99.99% allows about 52 minutes, and 99.999% allows about 5 minutes.
What Availability Actually Measures
Availability is the fraction of time a system is operational and reachable by its users, expressed as a percentage. If a service is up for 999 out of every 1000 minutes, it has 99.9% availability. People talk about it in "nines" because each extra nine is roughly a 10x reduction in allowed downtime, and each one is dramatically harder and more expensive to reach.
The math is simple. Availability equals uptime divided by total time. Over a 365 day year that is 525,600 minutes. So 99% lets you be down for about 3.65 days, 99.9% for about 8.77 hours, 99.99% for about 52.6 minutes, and 99.999% ("five nines") for about 5.26 minutes. A cloud provider's SLA of 99.95% means they promise no more than about 4.4 hours of downtime per year, or they owe you service credits.
Availability is not the same as reliability or durability. Reliability is about whether the system behaves correctly when it is up. Durability is about whether your stored data survives. A database can be 99.999% durable (it almost never loses data) but only 99.9% available (it sometimes refuses connections during a failover).
How You Actually Achieve It
The core lever is removing single points of failure through redundancy. Instead of one web server you run several behind a load balancer, so one crashing does not take the site down. Instead of one database you keep a replica that can be promoted if the primary dies. The system stays available as long as at least one healthy copy of each component exists.
Redundancy only helps if failures are detected fast and traffic is rerouted automatically. Health checks ping each instance every few seconds, and the load balancer stops sending traffic to anything that fails them. Failover promotes a standby to take over a dead primary. The time this takes, often called recovery time, directly eats into your availability budget.
There is a useful rule for combining components. When parts are in series, meaning all of them must work, you multiply their availabilities, so two 99.9% services chained together give only 99.8%. When parts are redundant, meaning any one is enough, the combined system is far more available than either alone. This is why a system with many dependencies needs each one to be very reliable, and why teams spread copies across separate availability zones and regions so one data center outage cannot take everything down.
When To Push For More Nines, And The Trade-offs
More availability is not free, and chasing nines you do not need wastes money and engineering time. The right target comes from the business. A payment system or hospital records system may justify five nines. An internal analytics dashboard is fine at three nines. The CAP theorem also forces a choice: during a network partition you can keep serving requests (stay available) or refuse them to stay strongly consistent, but not both.
Each extra nine roughly multiplies cost. Going from 99.9% to 99.99% often means multi-region deployment, automated failover, on-call rotations, chaos testing, and far more operational discipline. Going to five nines can mean active-active across regions with no manual steps anywhere in the recovery path, because even a few minutes of human reaction time blows the budget.
Most outages are not caused by hardware. They come from bad deploys, configuration mistakes, and dependency failures. So beyond redundancy, real availability work includes gradual rollouts, fast rollback, circuit breakers, rate limiting, and graceful degradation where the system serves a reduced experience instead of failing completely.
A Concrete Example
Consider an online store running in AWS. A single EC2 instance running both the app and the database might give you 99% availability, which is about 3.65 days of downtime a year and unacceptable for a shop.
Now spread it out. Run the app on three instances across three availability zones behind an Application Load Balancer, and use a managed database like RDS with a standby in a second zone for automatic failover. A single instance crash is invisible to customers because the load balancer reroutes traffic in seconds, and a database failure promotes the standby in a minute or two. This setup commonly reaches 99.95% to 99.99%.
If the business then needs to survive an entire AWS region going down, you replicate the whole stack into a second region with DNS-based or global load balancing that shifts traffic when the primary region fails health checks. That added complexity is what buys the jump toward five nines, and it is only worth it if a regional outage would genuinely cost more than the engineering and infrastructure it takes to defend against one.
Where it is used in production
Amazon Web Services
Publishes per-service SLAs such as 99.99% for EC2 across multiple availability zones and 99.999% for DynamoDB global tables, and pays service credits when it misses them.
Cloudflare
Runs an anycast network across hundreds of cities so a failed data center is simply routed around, keeping its DNS and CDN highly available.
Pioneered SRE practice built around error budgets, where the allowed downtime from an availability target is treated as a spendable resource that gates risky releases.
Netflix
Built Chaos Monkey to randomly kill production instances, proving the system stays available under failure rather than just hoping it will.
Frequently asked questions
- What is the difference between 99.9% and 99.99% availability?
- 99.9% (three nines) allows about 8.77 hours of downtime per year, while 99.99% (four nines) allows only about 52.6 minutes. Each added nine cuts allowed downtime by roughly 10x and usually requires multi-zone redundancy and automated failover.
- How is availability calculated?
- Availability equals uptime divided by total time, expressed as a percentage. For a year, total time is 525,600 minutes, so 99.99% availability means the system can be down for no more than about 52.6 of those minutes.
- Is availability the same as reliability?
- No. Availability is whether the system is up and reachable. Reliability is whether it behaves correctly when it is up. A service can be highly available but unreliable if it responds to every request with wrong results.
- Why is five nines so much harder than four nines?
- Five nines (99.999%) allows only about 5.26 minutes of downtime per year, which is less time than a human can react and fix something. It forces fully automated failover, active-active multi-region deployment, and removal of every manual step from the recovery path.
- Does adding more servers always increase availability?
- Only if they are truly redundant and failure is detected and rerouted automatically. If components are in series and all must work, adding more of them lowers combined availability because you multiply each part's availability together.
Learn Availability hands-on
This page explains the idea. The full lesson lets you step through the ring as servers join and leave, read the implementation, and check yourself with a quiz. It is one of 760+ lessons in the System Design Masterclass, from your first API call to distributed consensus. Eleven Foundation lessons are free, no signup. Lifetime access is ₹499 in India or $7.99 worldwide, one payment, no subscription.
Related lessons
Lessons that touch on Availability as part of a larger topic.
Availability Metrics
Measuring how reliably your system serves requests, the nines that define your reputation
intermediate · observability monitoring
CAP Theorem
Consistency, Availability, Partition Tolerance, you can only pick two, and you must pick partition tolerance
advanced · distributed systems core
Hinted Handoff
Store writes temporarily on behalf of unavailable nodes. Dynamo's approach to write availability
advanced · distributed systems core
Sloppy Quorum
Relaxed quorum rules that prioritize availability over strict replica placement
advanced · distributed systems core
Zero-Downtime Deployment
Deploying new versions without any period of unavailability for users
intermediate · devops cicd
See also
Related glossary terms you might want to look up next.
Fault Tolerance
A system's ability to keep operating correctly even when some of its components fail. Achieved through redundancy, replication, and graceful degradation.
CAP Theorem
In a distributed system, you can only guarantee two of three: Consistency, Availability, and Partition tolerance. You must choose your trade-off.
Redundancy
Duplicating critical components or functions so that if one fails, a backup takes over. The reason planes have two engines and databases have replicas.
Latency
The time delay between sending a request and getting a response. Amazon found every 100ms of extra latency costs 1% in sales.
Throughput
The number of operations a system can handle per unit of time. Think of it as how many cars a highway can move per hour.
Bandwidth
The maximum amount of data that can be transferred over a network in a given time. It's the width of the pipe, not how fast the water flows.