Architecture · · 2 min read
What 99.9% uptime really costs: lessons from a microservices platform at scale
Three nines sounds modest until you design for it. Error budgets, failure isolation, caching and team boundaries from a platform serving millions of users.
As a principal engineer I designed and implemented a microservices architecture that serves millions of users at 99.9% uptime, built on Spring Boot, Docker, Kubernetes and Redis. Three nines sounds modest. It’s about 43 minutes of downtime a month, total, across every deploy, dependency failure and bad config push. Here’s what it takes to live inside that budget.
Start with the error budget, not the architecture
99.9% isn’t an aspiration; it’s a budget. Around 43 minutes a month is what you get to spend on incidents, risky deploys and planned maintenance. Framing it this way changes conversations: when the budget is healthy, ship faster; when it’s burning, slow down and pay down reliability debt. It turns a vague argument between product and engineering into a number both sides agree on.
Microservices are an organisational decision
Service boundaries should follow team ownership and business capabilities, not technical layers. A service that three teams have to change together isn’t a microservice; it’s a distributed monolith with extra network hops. Before splitting anything I ask: which team owns this, can it deploy independently, and what happens to users when it’s down?
That last question drives the rest of the design.
Design for partial failure
In a distributed system something is always failing. The goal is that users barely notice.
- Timeouts everywhere. A missing timeout is how one slow dependency takes down everything upstream.
- Circuit breakers and bulkheads. Stop calling a failing dependency, and don’t let one misbehaving path exhaust shared thread pools or connections.
- Graceful degradation. Decide in advance what the product does without each dependency — stale data, a simpler experience, or a clear message — instead of a 500.
- Idempotent operations so retries are safe.
Cache deliberately
Redis did a lot of heavy lifting for read-heavy paths, but caching is a correctness decision disguised as a performance one. For every cached entity: what’s the acceptable staleness, how is it invalidated, and what happens when the cache is cold or down? A cache that the system can’t survive losing is a database you forgot to back up.
Make deploys boring
Most incidents are caused by change. Containerised services on Kubernetes gave us uniform, repeatable deploys with health checks, readiness probes and rolling updates. Pair that with progressive rollouts and fast rollback, and deploys stop being events. Boring deploys are a reliability feature.
Observe the user, not the server
CPU graphs don’t tell you whether users are happy. Service-level indicators should measure what users experience — request success rate, latency at the tail, key journey completion — and alerts should fire on symptoms users feel, not on every internal twitch. Fewer, better alerts keep on-call sustainable, and sustainable on-call is what keeps an uptime number real over years rather than quarters.
The takeaway for leaders
Reliability is a product feature with a price. Decide how much you want to pay for it, on purpose.
Three nines is achievable for most teams. Four nines is a different company. Knowing which one your business actually needs — and budgeting engineering time accordingly — is one of the highest-leverage calls a technology leader makes.