Alert fatigue kills incident response. When on-call engineers receive 200 alerts per shift, they stop reading them. Critical signals drown in a sea of warnings about disk usage on development servers, transient network blips that self-resolve, and threshold violations during expected maintenance windows. An effective alerting strategy ensures that every alert that reaches a human demands action and provides the context needed to take that action.
The Alert Fatigue Problem
Studies in healthcare — where alarm fatigue has been studied extensively — show that clinicians miss up to 85% of audible alarms when the false alarm rate exceeds 90%. The same dynamic applies to infrastructure alerting. When most alerts are false positives or require no action, engineers develop coping mechanisms: muting channels, ignoring pager notifications, and treating all alerts as low priority.
The cost extends beyond missed incidents. Teams with high alert volumes spend disproportionate time triaging and dismissing alerts instead of improving systems. Each false alarm interrupts deep work, taking 15-23 minutes to regain focus. Chronic alert fatigue leads to burnout, rotation avoidance, and eventual attrition of experienced on-call engineers.
Measuring your current alert health quantifies the problem. Track these metrics over rolling 30-day windows:
- Alert volume per on-call shift — high performers maintain under 2 actionable alerts per shift
- Signal-to-noise ratio — the percentage of alerts that required human intervention. Target above 80%
- Mean time to acknowledge (MTTA) — rising MTTA signals that engineers are losing trust in the alerting system
- Auto-resolved percentage — alerts that resolve before anyone acknowledges them. These are candidates for threshold adjustment or removal
Severity Classification Framework
Not every problem deserves the same response urgency. A severity framework maps the business impact of a condition to the response speed and escalation path it triggers.
Defining Severity by Business Impact
Severity levels should map to business outcomes, not technical metrics. "CPU above 90%" is a technical observation. "Checkout latency exceeds 3 seconds, affecting conversion rate" is a business impact that justifies a SEV 2 page. This reframing ensures that alerting priorities align with what actually matters to the organization.
SLO-Based Alerting
Service Level Objective (SLO) based alerting represents a fundamental shift from threshold monitoring to budget monitoring. Instead of alerting when a metric crosses a static threshold, SLO-based alerting fires when the rate of error budget consumption threatens to exhaust the budget before the SLO window closes.
Error Budget Burn Rate
An SLO of 99.9% availability over a 30-day window permits 43.2 minutes of downtime. The error budget burn rate measures how fast the budget is being consumed relative to the expected rate. A burn rate of 1.0 means the budget will exhaust exactly at the end of the window. A burn rate of 14.4 means the budget will exhaust in 1 hour — a situation demanding immediate response.
# Prometheus recording rules for SLO burn rate
# 99.9% availability SLO over 30 days
# Error ratio over 5m and 1h windows
- record: slo:error_ratio:5m
expr: |
sum(rate(http_requests_total{code=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
- record: slo:error_ratio:1h
expr: |
sum(rate(http_requests_total{code=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
# Burn rate (how fast we consume budget)
# burn_rate = error_ratio / (1 - SLO)
# For 99.9% SLO: (1 - 0.999) = 0.001
- record: slo:burn_rate:5m
expr: slo:error_ratio:5m / 0.001
- record: slo:burn_rate:1h
expr: slo:error_ratio:1h / 0.001
Multi-Window Alerting
Google's Site Reliability Engineering book popularized the multi-window, multi-burn-rate alerting pattern. The approach uses two time windows for each alert: a short window detects sudden spikes, and a long window confirms that the condition is sustained rather than transient.
A typical configuration:
- Page (SEV 1) — burn rate > 14.4x over 5 minutes AND burn rate > 14.4x over 1 hour. Detects severe incidents that will exhaust the budget in under 2 hours.
- Page (SEV 2) — burn rate > 6x over 30 minutes AND burn rate > 6x over 6 hours. Catches sustained degradation that will exhaust the budget within a day.
- Ticket (SEV 3) — burn rate > 3x over 2 hours AND burn rate > 1x over 3 days. Identifies slow burns that erode the error budget gradually.
Escalation Policy Design
An escalation policy defines what happens when an alert is not acknowledged or resolved within expected timeframes. The policy prevents critical alerts from being missed when the primary on-call engineer is unavailable, asleep through their phone, or overwhelmed with multiple simultaneous incidents.
Escalation Tiers
Structure escalation in three tiers:
- Primary on-call — receives the initial page. Given 5 minutes to acknowledge and 30 minutes to mitigate or escalate.
- Secondary on-call — receives the alert if the primary does not acknowledge. Often a different team member in a different timezone for follow-the-sun coverage.
- Engineering manager / incident commander — receives the alert if neither primary nor secondary acknowledges. At this tier, the response shifts from individual troubleshooting to organizational incident response.
Notification Channel Strategy
Different severity levels should use different notification channels to prevent desensitization. SEV 1 alerts use phone calls and SMS to cut through notification noise. SEV 2 alerts use push notifications from the incident management app. SEV 3 alerts post to Slack channels. SEV 4 information routes to dashboards and weekly summary emails.
Grouping related alerts prevents notification storms during cascading failures. When a database becomes unreachable, every service that depends on it will generate connection errors. Without grouping, the on-call engineer receives separate alerts for each service — dozens of pages for a single root cause. Alertmanager's grouping feature combines alerts sharing common labels (cluster, namespace, alertname) into a single notification.
Writing Actionable Runbooks
Every alert that can page a human must link to a runbook. A runbook transforms an alert from "something is wrong" into "here is how to diagnose and fix it." Without runbooks, incident response depends on tribal knowledge — knowledge that is unavailable when the engineer who wrote the alert is on vacation.
Runbook Structure
An effective runbook follows a consistent structure that engineers can navigate quickly at 3 AM:
- Alert context — what this alert means in plain language, what resource or service is affected, and what the user impact is.
- Diagnostic steps — ordered commands and queries to determine the root cause. Start with the most likely cause based on historical data.
- Mitigation actions — immediate steps to restore service, even if the root cause remains unresolved. Failover commands, traffic rerouting, cache clearing, service restarts.
- Escalation criteria — when to escalate to a specialist team, and how to reach them.
- Post-incident — what to document and which follow-up actions to create.
## Alert: API Error Rate > 5% (SEV 2)
### What's Happening
The API service is returning HTTP 5xx errors at a rate
exceeding 5% of total requests, indicating application
or infrastructure failure.
### Diagnostic Steps
1. Check error distribution by endpoint:
`sum by (path)(rate(http_requests_total{code=~"5.."}[5m]))`
2. Check recent deployments:
`kubectl rollout history deployment/api-server -n production`
3. Check database connectivity:
`kubectl exec -it deploy/api-server -- pg_isready -h db-primary`
4. Check dependent service health:
`curl -s http://auth-service:8080/health | jq .`
### Immediate Mitigation
- If deployment-related: `kubectl rollout undo deployment/api-server`
- If database: failover to replica (see DB Runbook #DB-003)
- If auth-service: enable bypass cache (feature flag auth_cache_bypass)
### Escalation
- Database issues → DBA on-call (PagerDuty: team-database)
- Network issues → Platform team (PagerDuty: team-platform)
Maintenance Windows and Suppression
Planned maintenance — database migrations, infrastructure upgrades, certificate rotations — generates expected alert conditions that should not page engineers. Alerting systems must support scheduled suppression windows that silence specific alerts during known maintenance periods.
Implement maintenance windows as time-bounded silences in Alertmanager or your incident management tool. Always scope silences narrowly: silence the specific alerts affected by maintenance on the specific hosts being maintained, not all alerts cluster-wide. Broad silences during maintenance windows have caused missed incidents when unrelated failures occur simultaneously.
Automatic suppression based on known conditions prevents another class of unnecessary pages. When a Kubernetes node enters NotReady state, suppress pod-level alerts for workloads on that node since they are consequences of the node failure, not independent problems. When a deployment rollout is in progress, suppress brief error rate spikes from pods that are cycling. These dependency-aware suppressions require encoding infrastructure topology into the alerting configuration.
On-Call Best Practices
The alerting strategy and the on-call practice that receives its output must be designed together. An excellent alerting configuration paired with a dysfunctional on-call rotation still results in poor incident response.
Rotation Design
Rotate on-call responsibility weekly to balance the burden while maintaining context. Shorter rotations (daily) prevent engineers from settling into the operational rhythm. Longer rotations (bi-weekly, monthly) risk burnout and create handoff gaps where context is lost.
Overlap the incoming and outgoing on-call engineers by 30-60 minutes to transfer active incidents, ongoing investigations, and pending maintenance. Document the shift handoff in a shared channel — what alerts fired, what was resolved, what remains open. This handoff log becomes historical context that log analysis alone cannot provide.
Continuous Improvement
Review every alert weekly. For each alert that fired, ask: Was it actionable? Did the engineer need to do something? Did the runbook help? Could the underlying condition have been prevented? Alerts that consistently auto-resolve, fire during expected operations, or lack useful runbooks are candidates for threshold adjustment, suppression rules, or removal.
Track the alert-to-incident ratio over time. A healthy system converts fewer than 20% of alerts into declared incidents. If the ratio exceeds 50%, the alerting system is either too sensitive (firing on conditions that are not actually problems) or too coarse (every alert represents a genuine failure, indicating broader reliability issues).