On-Call Practices for Performance Engineers
An on-call engineer who receives 15 pages during a single shift and resolves each one independently is not doing on-call well. They are firefighting. Sustainable on-call practice treats each page as both an immediate problem to solve and a signal about systemic reliability. The goal is not to respond faster but to be paged less often.
Performance-focused on-call presents unique challenges. Unlike availability alerts (service up or down), performance degradation exists on a continuum. A 20% increase in P99 latency may not be an incident, but a 200% increase certainly is. Defining when performance degradation crosses the threshold from "acceptable variation" to "actionable incident" requires clear SLO definitions and well-calibrated alerts.
Rotation Design
Team Size and Coverage
An on-call rotation needs at least 6 people for sustainable coverage. With fewer, each person is on-call too frequently. Weekly rotations with 6 engineers mean each person is primary on-call roughly 8-9 weeks per year—manageable. With 4 engineers, that becomes 13 weeks per year, approaching the burnout threshold.
| Rotation Size | Weeks On-Call/Year | Sustainability | Notes |
|---|---|---|---|
| 4 engineers | 13 | Unsustainable | High burnout risk, no slack for PTO |
| 6 engineers | 8-9 | Minimum viable | Works if page volume is low (< 5/week) |
| 8 engineers | 6-7 | Healthy | Handles moderate page volume |
| 10+ engineers | 5 | Comfortable | Room for high-complexity services |
Primary and Secondary
Assign both a primary and secondary on-call engineer. The primary handles all pages. The secondary takes over if the primary is unreachable or if the incident requires escalation. The secondary should be available within 15 minutes, not actively monitoring. This prevents both engineers from experiencing the full burden of on-call simultaneously.
Follow-the-Sun
For global teams, follow-the-sun rotations assign on-call coverage based on local business hours. Engineers in the Americas cover 08:00-16:00 UTC, Europe covers 08:00-16:00 CET, and Asia-Pacific covers remaining hours. Each engineer is only on-call during their daytime, eliminating overnight pages entirely. This requires at least three geographic regions with qualified on-call engineers in each.
Runbooks
A runbook is a documented procedure for responding to a specific alert condition. Without runbooks, on-call effectiveness depends entirely on individual expertise. With runbooks, a junior engineer can handle common performance incidents as effectively as a senior engineer.
Runbook Structure
Every runbook should follow a consistent structure:
- Alert description: What triggered this page? What metric crossed what threshold?
- Impact assessment: What is the user-facing impact? Which services are affected? What is the severity?
- Investigation steps: A prioritized list of dashboards to check, queries to run, and things to look for.
- Mitigation actions: Specific commands or procedures to reduce impact immediately (restart, rollback, scale up, feature flag toggle).
- Escalation criteria: When to escalate and to whom. "If mitigation does not resolve within 15 minutes, escalate to the database team."
- Root cause follow-up: Links to relevant documentation and instructions for filing the postmortem.
Performance-Specific Runbooks
Performance on-call requires runbooks for common degradation scenarios:
- Latency spike: Check recent deployments, database query plans, cache hit rates, connection pool utilization, and upstream service health
- Error rate increase: Identify affected endpoints, check deployment timeline, verify configuration changes, inspect error logs for patterns
- Resource exhaustion: Identify the exhausted resource (memory, connections, file descriptors), check for leaks, apply immediate mitigation (restart, scale), investigate root cause
- SLO burn rate alert: Assess remaining error budget, determine if the burn rate is accelerating or decelerating, decide on intervention vs. observation
Escalation Policies
Tiered Escalation
Define explicit escalation paths based on incident severity and duration:
| Time | SEV-1 (Outage) | SEV-2 (Degradation) | SEV-3 (Warning) |
|---|---|---|---|
| 0 min | Page primary on-call | Page primary on-call | Slack notification |
| 5 min | Page secondary + eng manager | Auto-acknowledge timer | — |
| 15 min | Incident commander declared | Page secondary if unack'd | Page primary if ongoing |
| 30 min | VP Engineering notified | Incident commander if unresolved | Secondary review |
| 60 min | Status page update | Postmortem trigger | — |
Cross-Team Escalation
Performance issues frequently span service boundaries. The API team's latency spike originates in the database team's slow query, which was caused by the data team's new ETL job. Escalation policies must include paths to adjacent teams with clear ownership boundaries: "If the investigation points to database layer issues, engage the database on-call via PagerDuty service X."
Alert Quality
The Page Ratio
Track the ratio of pages that require human intervention versus those that resolve automatically or require no action. A healthy target is at least 70% of pages requiring action. Below 50%, the on-call engineer is being woken up for nothing, and alert quality needs immediate improvement.
Performance Alert Calibration
Performance alerts need wider thresholds than availability alerts to account for normal variation. Latency varies with traffic patterns, time of day, and deployment cycles. An alert that fires on a 10% P99 increase will page the on-call engineer every time traffic spikes during business hours. Calibrate using SLO burn rate alerts instead of raw metric thresholds.
Preventing Burnout
On-call burnout does not come from incident volume alone. It comes from a combination of interrupt-driven work, sleep disruption, emotional weight of production responsibility, and lack of protected recovery time.
Compensation and Recovery
- Compensatory time off: An engineer who is paged overnight should receive comp time the following day. This is not optional—sleep-deprived engineers write bugs.
- On-call pay: Financial compensation acknowledges the personal cost of carrying a pager. Whether hourly, per-page, or as a shift differential, on-call should be compensated.
- Protected project time: On-call engineers should have reduced project work expectations. A team that expects full project velocity from someone carrying a pager will lose that person.
Toil Reduction
Every incident that pages the on-call engineer should generate a follow-up task: automate the response, fix the root cause, or improve the alert. Track the percentage of on-call time spent on toil (repetitive, automatable work) versus genuine engineering judgment. Chaos engineering exercises can proactively identify and address common failure modes before they generate pages.
Handoff Procedures
The transition between on-call shifts is a critical window where context gets lost. Formalize the handoff with a brief synchronous meeting (15 minutes maximum) covering:
- Active or recent incidents and their current status
- Known risks: upcoming deployments, maintenance windows, traffic events
- Action items from previous incidents that have not been resolved
- Any alerts that were acknowledged but not fully resolved
- Changes to runbooks or escalation paths since the last handoff
Key Takeaways
- Sustainable on-call requires at least 6 engineers per rotation. Primary and secondary roles prevent double burden. Follow-the-sun eliminates overnight pages.
- Runbooks are mandatory for every alert. Update them within 24 hours after any deviation during an incident.
- Define explicit escalation policies with time-based triggers and cross-team paths for multi-service performance issues.
- Track on-call health metrics: pages per shift, actionable ratio, MTTA, and overnight page frequency. Target > 70% actionable pages.
- Prevent burnout with comp time, on-call pay, reduced project expectations, and systematic toil reduction through automation.