Home › Observability & SRE › SLOs, SLIs, and Error Budgets

SLOs, SLIs, and Error Budgets: Measuring Reliability

Every engineering team faces the same tension: ship features faster or make the system more reliable. Without a framework to quantify reliability, this tension degenerates into political arguments between product managers who want velocity and operations engineers who want stability. Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets replace those arguments with data.

An SLO is a target for how reliable a service should be. An SLI is the metric that measures whether the service meets that target. An error budget is the amount of unreliability the target permits. Together, they create a decision framework: when the error budget has margin, ship features; when it is depleted, focus on reliability. This mechanism aligns incentives across the organization without requiring any team to win an argument.

Service Level Indicators: What to Measure

An SLI is a carefully defined metric that captures a dimension of user-perceived quality. The key phrase is user-perceived. Server CPU usage is not an SLI because users do not experience CPU usage. Request latency is an SLI because users directly experience slow responses.

Choosing the Right SLI Specification

Effective SLIs follow a ratio pattern: the proportion of events that are "good" divided by the total number of events. This produces a value between 0 and 1 (or 0% and 100%) that maps naturally to SLO targets.

SLI = (good events) / (total events) Availability SLI = (successful requests) / (total requests) Latency SLI = (requests faster than threshold) / (total requests) Quality SLI = (responses with correct data) / (total responses)

Common SLI categories for request-serving systems include:

SLI CategoryDefinitionMeasurement Point
AvailabilityProportion of requests that succeed (non-5xx)Load balancer access logs
LatencyProportion of requests served within thresholdApplication metrics, RUM data
CorrectnessProportion of responses returning valid dataEnd-to-end probes, data validation
FreshnessProportion of data updated within thresholdPipeline completion timestamps
ThroughputProportion of time the system serves above minimum capacityRequest counters, queue depth

SLI Implementation Principles

Where you measure the SLI dramatically affects its accuracy. Measuring availability at the application server misses failures caused by load balancers, DNS, or network issues. Measuring at the client captures all failure modes but introduces noise from client-side problems outside your control.

Best practice: Measure SLIs as close to the user as practical. For web services, the load balancer or API gateway is usually the best balance between accuracy and controllability. Supplement with Real User Monitoring data for end-to-end visibility.

Latency SLIs require particular care. A single latency threshold (e.g., "requests under 200ms") misses the distribution shape. A service where 95% of requests complete in 50ms but 5% take 10 seconds has a very different user experience than one where all requests take 190ms. Use multiple latency SLIs at different percentiles: p50, p95, and p99 each with their own thresholds and SLO targets.

Service Level Objectives: Setting Targets

An SLO is a target value for an SLI over a time window. "99.9% of requests succeed over a rolling 30-day window" is an SLO. The target is 99.9%, the SLI is availability (success ratio), and the compliance period is 30 days.

The Nines Table

Reliability targets are commonly expressed in "nines." Each additional nine dramatically reduces the permitted downtime:

SLO TargetPermitted Downtime/MonthPermitted Downtime/YearTypical Use Case
99% (two nines)7 hours 18 min3 days 15 hoursInternal tools, batch systems
99.5%3 hours 39 min1 day 19 hoursNon-critical APIs, staging
99.9% (three nines)43 minutes 50 sec8 hours 46 minMost production services
99.95%21 minutes 55 sec4 hours 23 minCustomer-facing platforms
99.99% (four nines)4 minutes 23 sec52 minutes 36 secPayment systems, auth services
99.999% (five nines)26 seconds5 minutes 15 secTelecom, emergency services

The jump from 99.9% to 99.99% is not a 0.09% improvement. It is a 10x reduction in permitted failure. The engineering investment required for each additional nine grows superlinearly: redundant infrastructure, automated failover, extensive testing, and reduced deployment velocity.

Setting Appropriate Targets

The most common mistake is setting SLOs too aggressively. An SLO of 99.99% availability sounds impressive but permits only 4 minutes of downtime per month. A single bad deployment that takes 10 minutes to roll back has already consumed two months of error budget. Teams with unrealistic SLOs either abandon them or become paralyzed by fear of deployment.

Start by measuring current performance. If your service currently achieves 99.7% availability, setting an SLO of 99.9% gives you a concrete improvement target. Setting 99.99% without the infrastructure to support it creates a target that generates noise (constant SLO violations) rather than signal.

SLO Target Selection Framework Step 1: Measure Current availability: 99.72% p99 latency: 380ms Error rate: 0.28% Step 2: Target SLO availability: 99.9% SLO latency p99: <300ms Improvement needed: 0.18% Step 3: Budget Error budget: 0.1% = 43 min downtime/month = ~2,592 failed requests/day Step 4: Policy Budget > 50% remaining → Ship features normally Budget < 50% → Increase testing, reduce deploy frequency | Budget depleted → Reliability sprint, freeze features

Error Budgets: The Decision Mechanism

An error budget is the inverse of an SLO. If the SLO is 99.9% availability, the error budget is 0.1%. Over a 30-day rolling window with 1 million requests per day, that error budget permits 30,000 failed requests. Every outage, every bug, every slow response consumes part of this budget.

How Error Budgets Work in Practice

Error budgets transform reliability from a vague aspiration into a concrete, measurable resource. The team starts each compliance period with a full budget. Incidents consume portions of it. The remaining budget determines what the team can do:

This mechanism makes the reliability vs. velocity tradeoff explicit and data-driven. Product managers cannot demand faster shipping when the error budget is depleted because the budget depletion proves the system cannot handle more change. Operations engineers cannot block deployments when the budget is healthy because the budget proves the system has margin for change.

Burn Rate Alerts

Traditional threshold alerts fire when a metric crosses a line: error rate exceeds 1%. Burn rate alerts are more sophisticated. They measure how fast the error budget is being consumed and alert when the current consumption rate would exhaust the budget before the compliance period ends.

burn_rate = (current_error_rate / error_budget_rate) # 30-day SLO window with 0.1% error budget # Normal burn rate = 1x (uses exactly 100% of budget in 30 days) # Alert thresholds: # 14.4x burn rate over 1 hour → budget exhausted in 2 days → PAGE # 6x burn rate over 6 hours → budget exhausted in 5 days → PAGE # 3x burn rate over 1 day → budget exhausted in 10 days → TICKET # 1x burn rate over 3 days → on track to exhaust budget → TICKET

Multi-window burn rate alerts reduce false positives by requiring elevated burn rates to persist across both short and long windows. A 14.4x burn rate sustained for only 2 minutes might be a transient spike. A 14.4x burn rate sustained for 1 hour with a 6-hour confirmation window is a real incident that demands immediate attention. This approach generates far fewer false alarms than simple threshold-based alerting strategies.

SLO Architecture Patterns

Request-Based vs. Window-Based SLOs

Request-based SLOs count individual events: "99.9% of requests succeed." Window-based SLOs count time intervals: "the system is available in 99.9% of 1-minute windows." The choice matters for bursty workloads.

Consider a service that receives 1,000 requests per minute. A 30-second outage affects 500 requests. Under a request-based SLO with a daily window, those 500 failures consume 500 out of a 1,440,000 daily budget (0.035%). Under a window-based SLO, that same outage consumes 1 out of 1,440 daily windows (0.07%). The request-based SLO is more forgiving for short outages affecting high-traffic periods because the denominator scales with traffic.

Composite SLOs

Complex services require multiple SLIs with separate SLO targets. An API might have:

Each SLO has its own error budget. The service is meeting its reliability contract only when all SLOs are within budget. A service that is fast but returns wrong data is not reliable, even if its availability and latency SLOs are green.

Dependency SLOs

A service cannot be more reliable than its least reliable critical dependency. If Service A depends on Service B, and Service B has a 99.9% availability SLO, then Service A cannot realistically target 99.99% availability without implementing fallback behavior for Service B's 0.1% failure rate.

Dependency analysis should inform SLO target setting. Map critical dependencies, understand their reliability characteristics, and set SLO targets that account for the cumulative failure probability of the dependency chain. Service-level monitoring provides the instrumentation needed to measure dependency reliability in practice rather than relying on published SLAs.

Organizational Implementation

Getting Buy-In

SLOs fail without organizational commitment. Engineering cannot unilaterally decide to freeze features when the error budget depletes. The policy must be agreed upon by engineering, product, and leadership before the first SLO is set. This pre-commitment is what gives error budgets their power: the decision is made rationally during calm times rather than emotionally during incidents.

SLO Review Cadence

SLOs are not permanent. Review them quarterly against three criteria:

  1. Are they too tight? If the error budget is consistently depleted, either the target is too aggressive or the system needs significant investment. Loosening the SLO honestly is better than abandoning the framework.
  2. Are they too loose? If the error budget is never consumed past 10%, the SLO is not providing useful signal. Tightening it creates room for the error budget policy to influence decisions.
  3. Do they capture user experience? If users complain despite green SLOs, the SLIs are measuring the wrong things. Add or revise SLIs to capture the failure modes users actually encounter.

Tracking error budget consumption over multiple compliance periods reveals patterns. A team that consistently depletes 60-80% of its budget is operating at a sustainable pace. A team that alternates between 0% and 100% consumption has reliability problems masked by quiet periods.

Error Budget Policies in Action

An error budget policy document codifies the responses to budget consumption levels. A practical policy includes:

Error Budget Policy: Order Processing Service SLO: 99.9% availability, 30-day rolling window Budget > 75%: - Standard deployment cadence (daily) - Feature work proceeds normally - Post-incident reviews within 3 business days Budget 50-75%: - Deployments require staged rollout (canary → 10% → 50% → 100%) - New features require load testing in staging - Post-incident reviews within 1 business day Budget 25-50%: - Deployments require explicit tech lead approval - 50% of sprint capacity allocated to reliability work - Automated rollback enabled for all deployments Budget < 25%: - Feature freeze: only reliability and bug fixes deploy - 100% of sprint capacity on reliability - Daily error budget review meetings Budget depleted: - Complete deployment freeze except critical security patches - Incident review for every budget-consuming event - Executive notification and recovery plan within 24 hours

The specificity matters. Vague policies ("we'll be more careful") provide no guidance. Specific policies ("staged rollouts required, 50% sprint capacity to reliability") create actionable expectations that every team member can follow without ambiguity.

Measuring SLOs with Existing Infrastructure

Implementing SLO measurement does not require specialized tooling. Most organizations can start with existing observability infrastructure:

The critical requirement is that SLI measurement is automated, continuous, and accessible to the entire team. A reliability metric that requires manual computation gets computed quarterly at best and provides no real-time guidance.

Key Takeaways