Home › Observability & SRE › Performance Regression Detection

Performance Regression Detection: Catching Slowdowns Before Users Do

A performance regression is a commit that makes things slower. It may add 15ms to a database query, increase memory allocation by 8%, or add one extra HTTP round trip to a user flow. Individually, each regression seems minor. But regressions accumulate. After six months of small regressions that nobody noticed, the application is 40% slower than the previous release and nobody knows which of the 2,000 merged commits is responsible.

Performance regression detection catches these slowdowns at the point they are introduced—during code review and CI/CD—before they compound into noticeable degradation.

Why Detection Is Hard

Unlike functional regressions, which produce objectively wrong output (a test either passes or fails), performance measurements are inherently noisy. The same code running on the same hardware produces different response times on each execution due to garbage collection, CPU cache behavior, OS scheduling, and background processes. A benchmark that shows 52ms on one run and 58ms on the next has not necessarily regressed—the 11% difference may be within normal variance.

This noise-to-signal problem means that naive comparison ("is the new result slower than the old result?") produces unacceptable false positive rates. Every other CI run flags a "regression" that is actually noise, and engineers learn to ignore the results. Effective detection requires statistical rigor.

Statistical Comparison Methods

Threshold-Based Detection

The simplest approach: flag a regression when the new result exceeds the baseline by more than a fixed percentage. "Alert if P99 latency increases by more than 10%." Simple to implement and understand, but the fixed threshold ignores the variance of the measurement. A 10% threshold is too loose for stable metrics with 1% variance and too tight for noisy metrics with 15% variance.

Standard Deviation Method

Compare the new measurement against the mean and standard deviation of the baseline measurements. Flag a regression when the new result exceeds mean + (k × standard deviation). With k=2, this flags results outside the 95th percentile of expected variation. With k=3, it flags outside the 99.7th percentile. This automatically adapts to metric variance but assumes a normal distribution, which performance data often is not.

Mann-Whitney U Test

A non-parametric statistical test that compares two samples without assuming normal distribution. Collect multiple measurements from both the baseline and candidate builds, then test whether the candidate sample is drawn from a distribution with a larger median. This is the most robust approach for performance data, which is typically right-skewed.

import numpy as np from scipy import stats def detect_regression(baseline_samples, candidate_samples, alpha=0.05, effect_size_threshold=0.05): """ Detect performance regression using Mann-Whitney U test. Returns (is_regression, p_value, effect_size) """ # Mann-Whitney U test (one-sided: is candidate slower?) statistic, p_value = stats.mannwhitneyu( baseline_samples, candidate_samples, alternative='less' # test if candidate > baseline ) # Effect size: relative difference in medians baseline_median = np.median(baseline_samples) candidate_median = np.median(candidate_samples) effect_size = (candidate_median - baseline_median) / baseline_median # Regression = statistically significant AND practically significant is_regression = (p_value < alpha) and (effect_size > effect_size_threshold) return is_regression, p_value, effect_size # Usage baseline = [48, 51, 49, 52, 50, 47, 53, 49, 51, 50] # ms candidate = [55, 58, 54, 57, 56, 53, 59, 55, 57, 56] # ms regressed, p, effect = detect_regression(baseline, candidate) # regressed=True, p=0.0001, effect=0.12 (12% slower)
Practical significance vs. statistical significance: A statistically significant result (low p-value) only means the difference is unlikely to be noise. It does not mean the difference matters. A 0.5% increase in latency may be statistically significant with enough samples but practically irrelevant. Always require both statistical significance and a minimum effect size before flagging a regression.

Baseline Management

The quality of regression detection depends entirely on the quality of the baseline. A baseline that is too old includes improvements that have since been made, making regressions harder to detect. A baseline that is too recent may have already been contaminated by recent regressions.

Rolling Baseline

Maintain a rolling baseline from the last N successful builds on the main branch. Each new merge updates the baseline with fresh measurements. This approach adapts naturally as the codebase evolves. The window size (N) balances stability (larger window = less noise) against sensitivity (smaller window = faster detection).

Tagged Baseline

Pin the baseline to a specific release tag. All performance comparisons are made against the last released version. This approach is simpler and more stable but may miss regressions that accumulate between releases. It works best for projects with frequent releases (weekly or biweekly).

Environment Consistency

Performance baselines are only valid if the measurement environment is consistent. A benchmark run on a shared CI runner with varying load is not comparable to one run on a dedicated machine. For reliable regression detection:

CI/CD Integration

Performance Gates

Add performance comparison as a required CI check. The pipeline runs benchmarks on the candidate branch, compares against the baseline, and reports results as a PR comment or CI status check. A detected regression blocks the merge until the author addresses it.

# GitHub Actions performance gate name: Performance Regression Check on: [pull_request] jobs: benchmark: runs-on: [self-hosted, perf-runner] steps: - uses: actions/checkout@v4 - name: Run benchmarks run: | # Run each benchmark 10 times for statistical power for i in $(seq 1 10); do ./run-benchmarks.sh >> candidate-results.json done - name: Fetch baseline run: | git checkout main for i in $(seq 1 10); do ./run-benchmarks.sh >> baseline-results.json done - name: Compare results run: | python compare-performance.py \ --baseline baseline-results.json \ --candidate candidate-results.json \ --threshold 0.05 \ --alpha 0.05 \ --output report.md - name: Comment on PR uses: actions/github-script@v7 with: script: | const fs = require('fs'); const report = fs.readFileSync('report.md', 'utf8'); github.rest.issues.createComment({ owner: context.repo.owner, repo: context.repo.repo, issue_number: context.issue.number, body: report });

Tiered Testing Strategy

Not all performance tests should run on every commit. Structure tests in tiers based on execution time and sensitivity:

TierWhenDurationScope
MicrobenchmarksEvery PR2-5 minHot functions, critical algorithms
Component benchmarksEvery merge to main10-20 minAPI endpoints, query performance
Full load testPre-release30-60 minComplete system under production load
Soak testWeekly4-8 hoursMemory leaks, resource exhaustion

What to Benchmark

Benchmark the operations that users experience. Not every function needs a benchmark. Focus on:

Handling False Positives

Even with statistical rigor, false positives will occur. The response to false positives determines whether the team trusts the regression detection system or ignores it:

Tracking Performance Over Time

Beyond per-commit regression detection, maintain a longitudinal view of performance across releases. Plot key metrics over time to detect gradual degradation that falls below per-commit detection thresholds. A 1% regression per week is undetectable on any single commit but adds up to 50% degradation over a year.

Build a performance trend dashboard that shows:

Key Takeaways