Performance Testing in CI/CD: Automated Regression Detection
Performance regressions are silent. A code change that adds 150ms to the p95 response time of your checkout endpoint won't break any functional test, won't trigger any error alert, and won't be visible in code review. It will only become apparent weeks later when enough small regressions accumulate into a noticeable user experience degradation — or when a traffic spike pushes the now-slower system past its capacity limit.
The only reliable way to catch performance regressions is to measure performance on every code change and compare against a baseline. This means integrating performance tests into your CI/CD pipeline: automated, reproducible, and blocking when thresholds are violated.
CI/CD Performance Testing Strategy
Not every performance test belongs in every pipeline stage. Full load tests that run for 30 minutes with hundreds of virtual users are valuable but too slow for every pull request. The strategy is to layer tests by speed and scope:
Tier 1: Every PR (1-3 minutes)
These tests run on every pull request and must complete fast enough to not slow development. Include:
- Unit benchmarks: Micro-benchmarks for critical algorithms, serialization, data processing functions. Detect algorithmic regressions (O(n) becoming O(n²)) before they reach integration.
- Lighthouse CI: Run Lighthouse against built pages to catch Core Web Vitals regressions from frontend changes — large image additions, excessive JavaScript, render-blocking resources.
- Bundle size checks: Track JavaScript and CSS bundle sizes. Alert when a bundle grows beyond a threshold. A dependency addition that adds 200KB to the bundle directly impacts LCP.
Tier 2: Per Merge to Main (5-10 minutes)
After a PR is approved and before or immediately after merge, run a quick API performance smoke test. This is a lightweight load test: 10-20 virtual users for 5 minutes against the critical path endpoints (login, search, checkout). The goal is not to find capacity limits but to detect response time regressions that indicate new N+1 queries, missing cache hits, or added middleware overhead.
// k6 CI smoke test — runs in 5 minutes
export const options = {
vus: 15,
duration: '5m',
thresholds: {
// Gate: fail the build if thresholds are violated
'http_req_duration{endpoint:checkout}': ['p(95)<500'],
'http_req_duration{endpoint:search}': ['p(95)<300'],
'http_req_duration{endpoint:homepage}': ['p(95)<200'],
http_req_failed: ['rate<0.01'], // Less than 1% errors
},
};
import http from 'k6/http';
import { sleep } from 'k6';
export default function() {
// Tag requests for per-endpoint thresholds
http.get('https://staging.example.com/', {
tags: { endpoint: 'homepage' }
});
sleep(1);
http.get('https://staging.example.com/search?q=test', {
tags: { endpoint: 'search' }
});
sleep(1);
http.post('https://staging.example.com/api/checkout',
JSON.stringify({ items: [{ id: 1, qty: 1 }] }),
{ headers: { 'Content-Type': 'application/json' },
tags: { endpoint: 'checkout' }
}
);
sleep(1);
}
Tier 3: Weekly or Pre-Release (30-60 minutes)
Full load tests with production-representative workloads, proper ramp-up patterns, and extended steady-state periods. Run these on a schedule (nightly or weekly) and before major releases. These tests validate capacity, not just per-request latency, and should run against a staging environment that mirrors production infrastructure.
Setting Performance Thresholds
Thresholds that are too tight will block deployments for normal variance and train the team to ignore performance gates. Thresholds that are too loose will miss real regressions. The right threshold balances sensitivity with stability.
Absolute vs. Relative Thresholds
| Approach | Definition | Pros | Cons |
|---|---|---|---|
| Absolute | p95 < 500ms | Simple, clear, aligns with SLOs | Doesn't detect gradual degradation within the threshold |
| Relative | < 10% slower than baseline | Catches all regressions regardless of absolute value | Requires baseline management, can false-positive from noise |
| Combined | Absolute gate + relative alert | Hard gate for SLOs, soft alert for trends | More complex to configure and maintain |
Handling Variance
CI environments introduce variance from shared infrastructure (other builds competing for CPU), cold starts (application and cache warmth), and non-deterministic workloads. To reduce false positives:
- Run tests multiple times: Take the median of 3 runs to smooth out outliers from infrastructure variance.
- Use dedicated test infrastructure: Dedicated containers or VMs for performance tests eliminate noisy neighbor effects.
- Warm up before measuring: Run 30-60 seconds of traffic before starting measurement to prime JIT compilation, caches, and connection pools.
- Track variance over time: If a test regularly produces p95 values between 180ms and 220ms, a threshold of 200ms will fail 50% of the time. Set the threshold above the observed variance band.
Pipeline Integration Patterns
GitHub Actions Example
# .github/workflows/perf-test.yml
name: Performance Smoke Test
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
lighthouse:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Build
run: npm run build
- name: Lighthouse CI
uses: treosh/lighthouse-ci-action@v11
with:
configPath: ./lighthouserc.json
uploadArtifacts: true
temporaryPublicStorage: true
api-perf:
runs-on: ubuntu-latest
if: github.event_name == 'push' # Only on merge to main
needs: [build-and-deploy-staging]
steps:
- uses: actions/checkout@v4
- name: Install k6
run: |
sudo gpg -k
sudo gpg --no-default-keyring --keyring /usr/share/keyrings/k6-archive-keyring.gpg --keyserver hkp://keyserver.ubuntu.com:80 --recv-keys C5AD17C747E3415A3642D57D77C6C491D6AC1D68
echo "deb [signed-by=/usr/share/keyrings/k6-archive-keyring.gpg] https://dl.k6.io/deb stable main" | sudo tee /etc/apt/sources.list.d/k6.list
sudo apt-get update
sudo apt-get install k6
- name: Run smoke test
run: k6 run tests/perf/smoke.js
env:
K6_CLOUD_TOKEN: ${{ secrets.K6_CLOUD_TOKEN }}
- name: Upload results
if: always()
uses: actions/upload-artifact@v4
with:
name: k6-results
path: results/
Lighthouse CI Configuration
// lighthouserc.json
{
"ci": {
"collect": {
"numberOfRuns": 3,
"startServerCommand": "npm run serve",
"url": [
"http://localhost:3000/",
"http://localhost:3000/search",
"http://localhost:3000/product/1"
]
},
"assert": {
"assertions": {
"categories:performance": ["error", { "minScore": 0.85 }],
"first-contentful-paint": ["warn", { "maxNumericValue": 2000 }],
"largest-contentful-paint": ["error", { "maxNumericValue": 2500 }],
"cumulative-layout-shift": ["error", { "maxNumericValue": 0.1 }],
"total-blocking-time": ["warn", { "maxNumericValue": 300 }],
"resource-summary:script:size": ["warn", { "maxNumericValue": 300000 }]
}
}
}
}
Baseline Management
For relative threshold comparison, you need a system to store and retrieve baseline performance data. The baseline should be updated after each successful test on the main branch, creating a rolling reference point.
Baseline Storage Options
- Git artifact: Store baseline JSON in the repository. Simple, versioned, but creates noise in commit history.
- CI artifact storage: GitHub Actions artifacts, GitLab CI artifacts. Tied to the CI system but easy to query.
- Dedicated service: k6 Cloud, Grafana Cloud, or a custom time-series database. Best for trend analysis and dashboards but adds infrastructure.
// Baseline comparison script
const fs = require('fs');
const baseline = JSON.parse(fs.readFileSync('baseline.json'));
const current = JSON.parse(fs.readFileSync('results.json'));
const regressions = [];
for (const [endpoint, metrics] of Object.entries(current)) {
const base = baseline[endpoint];
if (!base) continue;
const p95Change = (metrics.p95 - base.p95) / base.p95;
if (p95Change > 0.10) { // More than 10% slower
regressions.push({
endpoint,
baseline_p95: base.p95,
current_p95: metrics.p95,
change: `+${(p95Change * 100).toFixed(1)}%`
});
}
}
if (regressions.length > 0) {
console.error('Performance regressions detected:');
console.table(regressions);
process.exit(1); // Fail the pipeline
}
Common Implementation Mistakes
Testing Against Different Environments
A performance test against a local Docker environment tells you almost nothing about production performance. The database is empty, the CPU is shared with other build processes, and there are no network hops. If you cannot test against production-equivalent infrastructure, at least document the known differences and adjust thresholds accordingly.
No Warm-Up Phase
The first requests to a cold application include JIT compilation, cache priming, and connection establishment. If your test includes these cold-start requests in the measurement, the p95 will be artificially high and variance will be excessive. Always include a warm-up phase that generates traffic but does not record metrics.
Ignoring Infrastructure as Code
Performance test scripts, threshold configurations, and baseline data should be versioned alongside the application code. When a new feature intentionally changes performance characteristics (a new middleware, a more complex query), the performance thresholds should be updated in the same PR. Otherwise, the performance gate will block the deployment and someone will skip it, training the team to ignore performance gates.
Key Takeaways
- Layer performance tests by speed: unit benchmarks and Lighthouse on every PR (1-3 min), API smoke tests per merge (5-10 min), full load tests weekly or pre-release (30-60 min).
- Set CI thresholds at 80% of your SLO values to leave a buffer for production variance that controlled environments don't simulate.
- Reduce false positives: run multiple iterations, use dedicated infrastructure, warm up before measuring, and set thresholds above the observed variance band.
- Track baselines as rolling references: update after each successful main-branch test, compare PR results against the latest baseline.
- Version performance test scripts and thresholds alongside application code. Update thresholds in the same PR as intentional performance changes.
- A performance gate that gets routinely skipped is worse than no gate at all — it creates false confidence. Calibrate thresholds to minimize false positives.