SLOs, SLIs, and Error Budgets: Measuring Reliability
Every engineering team faces the same tension: ship features faster or make the system more reliable. Without a framework to quantify reliability, this tension degenerates into political arguments between product managers who want velocity and operations engineers who want stability. Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets replace those arguments with data.
An SLO is a target for how reliable a service should be. An SLI is the metric that measures whether the service meets that target. An error budget is the amount of unreliability the target permits. Together, they create a decision framework: when the error budget has margin, ship features; when it is depleted, focus on reliability. This mechanism aligns incentives across the organization without requiring any team to win an argument.
Service Level Indicators: What to Measure
An SLI is a carefully defined metric that captures a dimension of user-perceived quality. The key phrase is user-perceived. Server CPU usage is not an SLI because users do not experience CPU usage. Request latency is an SLI because users directly experience slow responses.
Choosing the Right SLI Specification
Effective SLIs follow a ratio pattern: the proportion of events that are "good" divided by the total number of events. This produces a value between 0 and 1 (or 0% and 100%) that maps naturally to SLO targets.
Common SLI categories for request-serving systems include:
| SLI Category | Definition | Measurement Point |
|---|---|---|
| Availability | Proportion of requests that succeed (non-5xx) | Load balancer access logs |
| Latency | Proportion of requests served within threshold | Application metrics, RUM data |
| Correctness | Proportion of responses returning valid data | End-to-end probes, data validation |
| Freshness | Proportion of data updated within threshold | Pipeline completion timestamps |
| Throughput | Proportion of time the system serves above minimum capacity | Request counters, queue depth |
SLI Implementation Principles
Where you measure the SLI dramatically affects its accuracy. Measuring availability at the application server misses failures caused by load balancers, DNS, or network issues. Measuring at the client captures all failure modes but introduces noise from client-side problems outside your control.
Latency SLIs require particular care. A single latency threshold (e.g., "requests under 200ms") misses the distribution shape. A service where 95% of requests complete in 50ms but 5% take 10 seconds has a very different user experience than one where all requests take 190ms. Use multiple latency SLIs at different percentiles: p50, p95, and p99 each with their own thresholds and SLO targets.
Service Level Objectives: Setting Targets
An SLO is a target value for an SLI over a time window. "99.9% of requests succeed over a rolling 30-day window" is an SLO. The target is 99.9%, the SLI is availability (success ratio), and the compliance period is 30 days.
The Nines Table
Reliability targets are commonly expressed in "nines." Each additional nine dramatically reduces the permitted downtime:
| SLO Target | Permitted Downtime/Month | Permitted Downtime/Year | Typical Use Case |
|---|---|---|---|
| 99% (two nines) | 7 hours 18 min | 3 days 15 hours | Internal tools, batch systems |
| 99.5% | 3 hours 39 min | 1 day 19 hours | Non-critical APIs, staging |
| 99.9% (three nines) | 43 minutes 50 sec | 8 hours 46 min | Most production services |
| 99.95% | 21 minutes 55 sec | 4 hours 23 min | Customer-facing platforms |
| 99.99% (four nines) | 4 minutes 23 sec | 52 minutes 36 sec | Payment systems, auth services |
| 99.999% (five nines) | 26 seconds | 5 minutes 15 sec | Telecom, emergency services |
The jump from 99.9% to 99.99% is not a 0.09% improvement. It is a 10x reduction in permitted failure. The engineering investment required for each additional nine grows superlinearly: redundant infrastructure, automated failover, extensive testing, and reduced deployment velocity.
Setting Appropriate Targets
The most common mistake is setting SLOs too aggressively. An SLO of 99.99% availability sounds impressive but permits only 4 minutes of downtime per month. A single bad deployment that takes 10 minutes to roll back has already consumed two months of error budget. Teams with unrealistic SLOs either abandon them or become paralyzed by fear of deployment.
Start by measuring current performance. If your service currently achieves 99.7% availability, setting an SLO of 99.9% gives you a concrete improvement target. Setting 99.99% without the infrastructure to support it creates a target that generates noise (constant SLO violations) rather than signal.
Error Budgets: The Decision Mechanism
An error budget is the inverse of an SLO. If the SLO is 99.9% availability, the error budget is 0.1%. Over a 30-day rolling window with 1 million requests per day, that error budget permits 30,000 failed requests. Every outage, every bug, every slow response consumes part of this budget.
How Error Budgets Work in Practice
Error budgets transform reliability from a vague aspiration into a concrete, measurable resource. The team starts each compliance period with a full budget. Incidents consume portions of it. The remaining budget determines what the team can do:
- Budget is healthy (more than 50% remaining): Deploy normally. Run experiments. Accept the risk of trying new approaches.
- Budget is thinning (25-50% remaining): Increase testing rigor. Require staged rollouts. Add monitoring for new deployments.
- Budget is critical (under 25%): Slow deployment cadence. Require explicit approval for changes. Focus engineering time on reliability improvements.
- Budget is depleted: Freeze feature deployments. All engineering effort goes to reliability. Only emergency fixes ship.
This mechanism makes the reliability vs. velocity tradeoff explicit and data-driven. Product managers cannot demand faster shipping when the error budget is depleted because the budget depletion proves the system cannot handle more change. Operations engineers cannot block deployments when the budget is healthy because the budget proves the system has margin for change.
Burn Rate Alerts
Traditional threshold alerts fire when a metric crosses a line: error rate exceeds 1%. Burn rate alerts are more sophisticated. They measure how fast the error budget is being consumed and alert when the current consumption rate would exhaust the budget before the compliance period ends.
Multi-window burn rate alerts reduce false positives by requiring elevated burn rates to persist across both short and long windows. A 14.4x burn rate sustained for only 2 minutes might be a transient spike. A 14.4x burn rate sustained for 1 hour with a 6-hour confirmation window is a real incident that demands immediate attention. This approach generates far fewer false alarms than simple threshold-based alerting strategies.
SLO Architecture Patterns
Request-Based vs. Window-Based SLOs
Request-based SLOs count individual events: "99.9% of requests succeed." Window-based SLOs count time intervals: "the system is available in 99.9% of 1-minute windows." The choice matters for bursty workloads.
Consider a service that receives 1,000 requests per minute. A 30-second outage affects 500 requests. Under a request-based SLO with a daily window, those 500 failures consume 500 out of a 1,440,000 daily budget (0.035%). Under a window-based SLO, that same outage consumes 1 out of 1,440 daily windows (0.07%). The request-based SLO is more forgiving for short outages affecting high-traffic periods because the denominator scales with traffic.
Composite SLOs
Complex services require multiple SLIs with separate SLO targets. An API might have:
- Availability SLO: 99.95% of requests return non-5xx responses
- Latency SLO: 99% of requests complete within 200ms
- Latency SLO: 99.9% of requests complete within 1000ms
- Correctness SLO: 99.99% of responses contain valid data
Each SLO has its own error budget. The service is meeting its reliability contract only when all SLOs are within budget. A service that is fast but returns wrong data is not reliable, even if its availability and latency SLOs are green.
Dependency SLOs
A service cannot be more reliable than its least reliable critical dependency. If Service A depends on Service B, and Service B has a 99.9% availability SLO, then Service A cannot realistically target 99.99% availability without implementing fallback behavior for Service B's 0.1% failure rate.
Dependency analysis should inform SLO target setting. Map critical dependencies, understand their reliability characteristics, and set SLO targets that account for the cumulative failure probability of the dependency chain. Service-level monitoring provides the instrumentation needed to measure dependency reliability in practice rather than relying on published SLAs.
Organizational Implementation
Getting Buy-In
SLOs fail without organizational commitment. Engineering cannot unilaterally decide to freeze features when the error budget depletes. The policy must be agreed upon by engineering, product, and leadership before the first SLO is set. This pre-commitment is what gives error budgets their power: the decision is made rationally during calm times rather than emotionally during incidents.
SLO Review Cadence
SLOs are not permanent. Review them quarterly against three criteria:
- Are they too tight? If the error budget is consistently depleted, either the target is too aggressive or the system needs significant investment. Loosening the SLO honestly is better than abandoning the framework.
- Are they too loose? If the error budget is never consumed past 10%, the SLO is not providing useful signal. Tightening it creates room for the error budget policy to influence decisions.
- Do they capture user experience? If users complain despite green SLOs, the SLIs are measuring the wrong things. Add or revise SLIs to capture the failure modes users actually encounter.
Tracking error budget consumption over multiple compliance periods reveals patterns. A team that consistently depletes 60-80% of its budget is operating at a sustainable pace. A team that alternates between 0% and 100% consumption has reliability problems masked by quiet periods.
Error Budget Policies in Action
An error budget policy document codifies the responses to budget consumption levels. A practical policy includes:
The specificity matters. Vague policies ("we'll be more careful") provide no guidance. Specific policies ("staged rollouts required, 50% sprint capacity to reliability") create actionable expectations that every team member can follow without ambiguity.
Measuring SLOs with Existing Infrastructure
Implementing SLO measurement does not require specialized tooling. Most organizations can start with existing observability infrastructure:
- Load balancer logs: Parse access logs to compute availability SLIs (non-5xx response ratio) and latency SLIs (response time distribution).
- Application metrics: Use existing Prometheus, Datadog, or CloudWatch metrics to track request counts, error counts, and latency histograms.
- Synthetic monitoring: Run synthetic probes from external locations to measure availability from the user's perspective, catching failures that internal metrics miss.
The critical requirement is that SLI measurement is automated, continuous, and accessible to the entire team. A reliability metric that requires manual computation gets computed quarterly at best and provides no real-time guidance.
Key Takeaways
- SLIs measure user-perceived quality as a ratio of good events to total events. Measure as close to the user as practical.
- SLOs set reliability targets for SLIs over defined time windows. Start with current performance and set achievable improvement targets.
- Error budgets quantify permitted unreliability. They transform the reliability vs. velocity argument from politics into data.
- Burn rate alerts detect budget consumption patterns early, before the budget depletes, and generate fewer false positives than simple thresholds.
- Error budget policies must be agreed upon by engineering, product, and leadership before the first SLO is set. Specificity in the policy creates actionable expectations.