Home › Observability & SRE › Incident Response for Performance Degradation

Incident Response for Performance Degradation

Performance incidents are not binary. A server crash is obvious: the system is up or it is down. Performance degradation is subtle. Response times creep from 200ms to 800ms over two hours. Error rates climb from 0.1% to 2.3% during peak traffic. A memory leak slowly reduces available capacity until garbage collection pauses start affecting user experience. These gradual failures are harder to detect, harder to triage, and harder to resolve than outright crashes.

Effective incident response for performance issues requires a structured process that addresses four phases: detection (recognizing something is wrong), triage (understanding what is wrong and its severity), mitigation (stopping the bleeding), and resolution (fixing the root cause). Each phase has distinct objectives, tooling requirements, and communication protocols.

Phase 1: Detection

The fastest path to detection is automated alerting. The slowest is a customer complaint. The gap between these two paths directly measures the maturity of your monitoring infrastructure.

Detection Time Budget

Detection time is the interval between when a problem begins and when someone starts investigating. For performance degradation, this interval often dwarfs the actual fix time. A p99 latency regression from 200ms to 2 seconds might take 45 minutes to detect but only 5 minutes to fix with a rollback.

Three detection mechanisms, ordered by speed:

  1. Automated alerts (minutes): SLO burn rate alerts detect sustained degradation within minutes. They fire when error budget consumption exceeds the sustainable rate, catching both sudden spikes and gradual declines.
  2. Anomaly detection (5-15 minutes): Machine learning models that baseline normal behavior and flag statistical deviations. Effective for catching novel failure modes that predetermined thresholds miss.
  3. Human observation (15-60+ minutes): Engineers noticing problems during routine dashboard reviews, or users reporting issues through support channels. If this is your primary detection mechanism, your alerting needs work.
Incident Timeline: Where Time Goes Detect Triage Mitigate Resolve root cause Postmortem 5-45 min 10-30 min 5-15 min hours to days 1-5 days User impact starts User impact ends Total user impact = Detection + Triage + Mitigation time

The critical insight is that user impact equals detection time plus triage time plus mitigation time. Resolution time (fixing the root cause) happens after mitigation restores service, so it does not contribute to user-facing impact. This is why "roll back first, investigate later" is almost always the right strategy.

Phase 2: Triage

Triage answers three questions: What is broken? How bad is it? Who should fix it?

Severity Classification

SeverityImpact DescriptionResponseCommunication
SEV-1 (Critical)Complete outage or >50% of users affectedAll hands, war room, 15-min updatesStatus page, executive notification
SEV-2 (Major)Significant degradation, 10-50% of usersOn-call + backup, 30-min updatesStatus page, team notification
SEV-3 (Minor)Limited degradation, <10% of usersOn-call engineer, hourly updatesInternal team channel
SEV-4 (Low)Cosmetic or non-user-facing issueNormal priority ticketTeam standup mention

Performance incidents often start as SEV-3 and escalate. A 20% latency increase during off-peak hours is a SEV-3. The same increase during peak traffic affecting checkout flows is a SEV-1. Severity classification must account for context: time of day, affected services, and business impact.

Triage Checklist

When an alert fires, the on-call engineer should work through a structured checklist rather than relying on intuition:

  1. Verify the alert is real. Check the raw data behind the alert. Is this a genuine degradation or a monitoring artifact? False positive alerts waste response time and erode trust in the alerting system.
  2. Scope the impact. Which services are affected? Which users? Which regions? Use observability tooling to narrow from "something is slow" to "the inventory service p99 latency in us-east-1 is 5x normal."
  3. Correlate with changes. Check the deployment log. Was anything deployed in the last 2 hours? Configuration changes? Infrastructure modifications? Database migrations? The most common cause of performance degradation is a recent change.
  4. Check dependencies. Is the degradation caused by an upstream or downstream service? A slow database makes every service that queries it slow. Fixing the symptom in the wrong service wastes time.
  5. Classify severity. Based on the scope and impact, assign a severity level and invoke the corresponding response protocol.

Phase 3: Mitigation

Mitigation restores service quality. It is not the same as resolution. A rollback that reverts a bad deployment is mitigation: service is restored, but the underlying code change still needs to be fixed. The goal of mitigation is to minimize user impact as quickly as possible.

Mitigation Strategies

Ordered by speed and risk:

Anti-pattern: "Let me just fix it." Engineers under pressure often bypass mitigation and go straight to debugging the root cause. This extends user impact from minutes (rollback time) to hours (debugging time). Always mitigate first, investigate second.

Rollback Decision Framework

The decision to rollback should take less than 5 minutes. Use this framework:

IF deployment in last 2 hours AND performance correlates with deploy time: → ROLLBACK immediately → Investigate the deployment after service is restored IF no recent deployment AND issue is resource exhaustion: → SCALE horizontally while investigating → Check for memory leaks, connection pool exhaustion, runaway queries IF issue is intermittent AND affects specific users/regions: → Enable detailed tracing for affected traffic → Check CDN, DNS, and network path for the affected region IF rollback attempted AND issue persists: → Issue predates the deployment; stop rolling back further → Escalate severity and expand investigation scope

Phase 4: Resolution and Postmortem

Root Cause Analysis

After mitigation restores service, the team shifts to root cause analysis. This is methodical work that should not be rushed. Common performance root causes fall into categories:

Use the "five whys" technique to distinguish symptoms from causes. "The database is slow" is a symptom. "The database is slow because a missing index causes full table scans on a table that grew past 5 million rows after last week's data migration" is a root cause with an actionable fix.

Blameless Postmortems

A blameless postmortem focuses on systemic failures rather than individual mistakes. The premise is that the engineer who deployed the bad code was operating within a system that allowed that code to reach production. The fix is not "be more careful" but rather "add a performance test that catches this regression before deployment."

A postmortem document should include:

  1. Timeline: Precise chronology of events from root cause introduction through detection, triage, mitigation, and resolution.
  2. Impact: Quantified user impact: number of affected users, duration of degradation, error budget consumed, estimated revenue impact.
  3. Root cause: Technical explanation of what went wrong and why existing safeguards did not prevent it.
  4. What went well: Aspects of the response that worked effectively. Reinforcing good practices is as important as fixing gaps.
  5. Action items: Specific, assigned, time-bound improvements. "Improve monitoring" is not an action item. "Add latency SLI for payment-service p99 with 300ms threshold, assigned to Elena, due by July 28" is an action item.

Postmortem action items should address all four phases. Detection improvements reduce time-to-alert. Triage improvements include better runbooks. Mitigation improvements add rollback automation or feature flags. Prevention improvements include automated regression detection in the CI/CD pipeline and performance testing in CI/CD.

Communication During Incidents

Internal Communication

Incident communication should be structured and predictable. Designate an Incident Commander (IC) who owns communication and coordination, separate from the engineers doing the debugging. The IC posts regular updates to the incident channel at the cadence defined by the severity level.

Each update follows a standard format:

STATUS UPDATE - [SEV-2] Payment latency degradation Time: 2026-07-14 14:45 UTC (30 min into incident) Current state: Payment service p99 latency at 3.2s (normal: 200ms) Impact: ~15% of checkout attempts timing out Actions taken: Rolled back deploy #4521, latency improving Next update: 15:00 UTC or sooner if status changes IC: Carlos Reyes

External Communication

Status pages should reflect the user experience, not internal severity. Users do not care about SEV levels. They care about whether they can complete their tasks. Use language like "Some users may experience slower than normal page loads" rather than "SEV-2 incident affecting payment-service p99 latency."

Update the status page when the impact changes, not on a fixed cadence. An update that says "we're still investigating" provides no value and erodes trust. An update that says "we've identified the cause and expect resolution within 30 minutes" provides useful information.

Building Incident Readiness

Runbooks

Runbooks are step-by-step guides for common incident scenarios. A good runbook turns a 30-minute triage by an expert into a 10-minute triage by any on-call engineer. Every service should have runbooks for its top failure modes, covering the exact commands to run, metrics to check, and escalation contacts.

Game Days

Practicing incident response under controlled conditions builds muscle memory that carries over to real incidents. Chaos engineering exercises simulate performance degradation, forcing teams to practice detection, triage, and mitigation against a known problem. The learning comes from discovering gaps in alerting, runbooks, and communication that only surface under the pressure of an active incident.

On-Call Preparation

Effective on-call practices ensure that the responding engineer has the access, context, and tools to handle incidents independently. This means pre-provisioned access to dashboards, deployment systems, and communication channels. An engineer who spends the first 20 minutes of an incident requesting access to the metrics dashboard is an engineer who added 20 minutes to the incident duration.

Key Takeaways