Home › Observability & SRE › On-Call Practices for Performance

On-Call Practices for Performance Engineers

An on-call engineer who receives 15 pages during a single shift and resolves each one independently is not doing on-call well. They are firefighting. Sustainable on-call practice treats each page as both an immediate problem to solve and a signal about systemic reliability. The goal is not to respond faster but to be paged less often.

Performance-focused on-call presents unique challenges. Unlike availability alerts (service up or down), performance degradation exists on a continuum. A 20% increase in P99 latency may not be an incident, but a 200% increase certainly is. Defining when performance degradation crosses the threshold from "acceptable variation" to "actionable incident" requires clear SLO definitions and well-calibrated alerts.

Rotation Design

Team Size and Coverage

An on-call rotation needs at least 6 people for sustainable coverage. With fewer, each person is on-call too frequently. Weekly rotations with 6 engineers mean each person is primary on-call roughly 8-9 weeks per year—manageable. With 4 engineers, that becomes 13 weeks per year, approaching the burnout threshold.

Rotation SizeWeeks On-Call/YearSustainabilityNotes
4 engineers13UnsustainableHigh burnout risk, no slack for PTO
6 engineers8-9Minimum viableWorks if page volume is low (< 5/week)
8 engineers6-7HealthyHandles moderate page volume
10+ engineers5ComfortableRoom for high-complexity services

Primary and Secondary

Assign both a primary and secondary on-call engineer. The primary handles all pages. The secondary takes over if the primary is unreachable or if the incident requires escalation. The secondary should be available within 15 minutes, not actively monitoring. This prevents both engineers from experiencing the full burden of on-call simultaneously.

Follow-the-Sun

For global teams, follow-the-sun rotations assign on-call coverage based on local business hours. Engineers in the Americas cover 08:00-16:00 UTC, Europe covers 08:00-16:00 CET, and Asia-Pacific covers remaining hours. Each engineer is only on-call during their daytime, eliminating overnight pages entirely. This requires at least three geographic regions with qualified on-call engineers in each.

Runbooks

A runbook is a documented procedure for responding to a specific alert condition. Without runbooks, on-call effectiveness depends entirely on individual expertise. With runbooks, a junior engineer can handle common performance incidents as effectively as a senior engineer.

Runbook Structure

Every runbook should follow a consistent structure:

  1. Alert description: What triggered this page? What metric crossed what threshold?
  2. Impact assessment: What is the user-facing impact? Which services are affected? What is the severity?
  3. Investigation steps: A prioritized list of dashboards to check, queries to run, and things to look for.
  4. Mitigation actions: Specific commands or procedures to reduce impact immediately (restart, rollback, scale up, feature flag toggle).
  5. Escalation criteria: When to escalate and to whom. "If mitigation does not resolve within 15 minutes, escalate to the database team."
  6. Root cause follow-up: Links to relevant documentation and instructions for filing the postmortem.
Runbook maintenance rule: If you deviated from the runbook during an incident, update the runbook within 24 hours. Runbooks that do not reflect reality are worse than no runbooks because they waste time and erode trust in documentation.

Performance-Specific Runbooks

Performance on-call requires runbooks for common degradation scenarios:

Escalation Policies

Tiered Escalation

Define explicit escalation paths based on incident severity and duration:

TimeSEV-1 (Outage)SEV-2 (Degradation)SEV-3 (Warning)
0 minPage primary on-callPage primary on-callSlack notification
5 minPage secondary + eng managerAuto-acknowledge timer—
15 minIncident commander declaredPage secondary if unack'dPage primary if ongoing
30 minVP Engineering notifiedIncident commander if unresolvedSecondary review
60 minStatus page updatePostmortem trigger—

Cross-Team Escalation

Performance issues frequently span service boundaries. The API team's latency spike originates in the database team's slow query, which was caused by the data team's new ETL job. Escalation policies must include paths to adjacent teams with clear ownership boundaries: "If the investigation points to database layer issues, engage the database on-call via PagerDuty service X."

Alert Quality

The Page Ratio

Track the ratio of pages that require human intervention versus those that resolve automatically or require no action. A healthy target is at least 70% of pages requiring action. Below 50%, the on-call engineer is being woken up for nothing, and alert quality needs immediate improvement.

# On-call health metrics to track weekly Pages per shift: Target < 5 Actionable page ratio: Target > 70% Mean time to acknowledge: Target < 5 min Mean time to mitigate: Target < 30 min Overnight pages: Target < 1 per shift Unique alert types: Track for diversity Repeat pages (same alert): Flag for automation

Performance Alert Calibration

Performance alerts need wider thresholds than availability alerts to account for normal variation. Latency varies with traffic patterns, time of day, and deployment cycles. An alert that fires on a 10% P99 increase will page the on-call engineer every time traffic spikes during business hours. Calibrate using SLO burn rate alerts instead of raw metric thresholds.

Preventing Burnout

On-call burnout does not come from incident volume alone. It comes from a combination of interrupt-driven work, sleep disruption, emotional weight of production responsibility, and lack of protected recovery time.

Compensation and Recovery

Toil Reduction

Every incident that pages the on-call engineer should generate a follow-up task: automate the response, fix the root cause, or improve the alert. Track the percentage of on-call time spent on toil (repetitive, automatable work) versus genuine engineering judgment. Chaos engineering exercises can proactively identify and address common failure modes before they generate pages.

Handoff Procedures

The transition between on-call shifts is a critical window where context gets lost. Formalize the handoff with a brief synchronous meeting (15 minutes maximum) covering:

Key Takeaways