Stress Testing Web Applications: Finding Your System's Breaking Point
Stress testing pushes a system beyond its designed capacity to find where and how it fails. While load testing validates that the system handles expected traffic, stress testing deliberately exceeds that threshold. The goal is not to prevent failure — every system has limits — but to understand the failure mode. A system that degrades gracefully under overload (slower responses, queued requests, shed non-essential work) is fundamentally different from one that cascades into complete unavailability (connection pool exhaustion, out-of-memory crash, data corruption).
The distinction matters because production systems will experience overload. Traffic spikes from viral content, marketing campaigns, seasonal events, or DDoS attacks can exceed design capacity by 5x or more. Knowing whether your system sheds load gracefully or falls over entirely determines your incident response strategy.
Stress Test Design
A stress test uses the same workload model as a baseline load test but applies it at increasing intensity. The key design decision is how fast to increase load and at what increments.
Step-Load Approach
The most informative stress test uses a step-load pattern: increase concurrent users by a fixed increment (such as 10-20% of baseline load) every 5 minutes and hold steady at each step to allow the system to stabilize. This creates clear data points showing how each metric (response time, error rate, CPU, memory, connection count) changes at each load level.
// k6 step-load stress test configuration
export const options = {
stages: [
{ duration: '5m', target: 100 }, // Baseline: 100 users
{ duration: '5m', target: 100 }, // Hold at baseline
{ duration: '5m', target: 200 }, // Step 1: 2x baseline
{ duration: '5m', target: 200 }, // Hold
{ duration: '5m', target: 400 }, // Step 2: 4x baseline
{ duration: '5m', target: 400 }, // Hold
{ duration: '5m', target: 600 }, // Step 3: 6x baseline
{ duration: '5m', target: 600 }, // Hold
{ duration: '5m', target: 800 }, // Step 4: 8x baseline
{ duration: '5m', target: 800 }, // Hold
{ duration: '5m', target: 1000 }, // Step 5: 10x baseline
{ duration: '5m', target: 1000 }, // Hold
{ duration: '5m', target: 0 }, // Ramp down (recovery)
],
thresholds: {
http_req_duration: ['p(95)<2000'], // Track but don't abort
http_req_failed: ['rate<0.50'], // Abort if 50% errors
},
};
Failure Modes to Identify
Stress testing reveals specific failure modes, each with different impact and mitigation strategies:
| Failure Mode | Symptom | Root Cause | Severity |
|---|---|---|---|
| Connection pool exhaustion | Sudden spike in response time + connection timeout errors | All DB/HTTP connections in use, new requests queue indefinitely | High — cascading |
| Memory pressure | Gradual response time increase, OOM kill | Request buffers, session data, or caches grow without bound | Critical — process dies |
| Thread/worker starvation | New requests rejected or timeout while CPU is low | All workers blocked on slow I/O (DB queries, external APIs) | High — total stall |
| CPU saturation | Linear response time increase with load | Compute-bound processing (serialization, encryption, rendering) | Medium — degrades gracefully |
| Disk I/O saturation | Response time spikes correlated with iowait | Heavy logging, swap usage, or DB writes exceed disk throughput | Medium — varies by storage |
| Cascading failure | One service fails, then upstream services queue and fail | No circuit breakers, unbounded retry, tight coupling | Critical — full outage |
Connection Pool Exhaustion
This is the most common failure mode in web applications under stress. A database connection pool with 20 connections works fine at normal load. Under 5x load, all 20 connections are busy with active queries. New requests wait in the connection pool queue. If the queue has no timeout, requests wait indefinitely — the application appears hung rather than returning an error. If the queue has a short timeout, requests fail fast and clients can retry or receive an error page, which is the better failure mode.
// PostgreSQL connection pool with bounded waiting
const pool = new Pool({
max: 20, // Maximum connections
connectionTimeoutMillis: 5000, // Fail fast if no connection available
idleTimeoutMillis: 30000, // Release idle connections
// Under stress, new requests get a connection timeout
// error after 5 seconds rather than waiting indefinitely
});
Cascading Failures
Cascading failure occurs when overload in one component propagates to dependent components. Service A calls Service B. Under stress, Service B slows down. Service A's requests to Service B accumulate, consuming Service A's threads and connections. Service A can no longer handle its own requests, even those that don't depend on Service B. Now Service C, which depends on Service A, starts failing too.
Circuit breakers prevent cascading failures by detecting when a downstream service is failing and short-circuiting requests to it. This lets the calling service fail fast on the specific functionality that depends on the failing service while continuing to serve other requests normally.
Recovery Testing
How quickly the system recovers after overload is removed is as important as how it behaves under overload. Some systems recover instantly when excess load is removed. Others enter a "death spiral" where recovery is impossible without manual intervention (restarting processes, clearing queues, dropping connections).
Recovery Patterns
- Clean recovery: Response time and error rate return to baseline within 60 seconds of load removal. Connection pools drain, queues empty, caches repopulate. No intervention needed.
- Slow recovery: System gradually returns to normal over 5-10 minutes. Often caused by cache repopulation (cold cache after eviction under memory pressure) or connection pool rebalancing.
- Stuck state: System does not recover without intervention. Typically caused by leaked resources (connections not returned to pool), deadlocked threads, or corrupted state. Requires process restart.
- Thundering herd: Queued requests all execute simultaneously when capacity returns, immediately re-overloading the system. Requires request shedding or gradual queue drain.
Overload Protection Mechanisms
Rate Limiting
Rate limiting caps the number of requests accepted per unit of time. Requests that exceed the limit receive an immediate 429 (Too Many Requests) response. This protects backend resources by ensuring the system never processes more than its tested capacity. The key decision is where to apply rate limiting: at the load balancer (broadest protection), the API gateway (per-route granularity), or within the application (per-user fairness).
Load Shedding
Load shedding selectively drops requests when the system is overloaded, prioritizing important requests over less important ones. A checkout endpoint that generates revenue should be protected even when product browsing endpoints are overwhelmed. Implement load shedding using queue depth or response time as the trigger, and priority classification based on endpoint, user type, or request source.
Backpressure
Backpressure signals upstream components to slow down when downstream capacity is reached. Instead of accepting requests and queuing them (which eventually leads to memory exhaustion), the system signals that it cannot accept more work. HTTP 503 (Service Unavailable) with a Retry-After header is the standard backpressure signal for HTTP APIs.
Monitoring During Stress Tests
Run stress tests with comprehensive infrastructure monitoring and observability in place. The test tool shows client-side symptoms (response time degradation, errors), but monitoring shows the underlying cause. Correlating the two reveals which resource bottleneck causes each symptom.
Key monitoring views during stress testing:
- CPU per core: Total CPU might look fine at 60% while one core is saturated at 100% (common with single-threaded operations).
- Memory breakdown: Heap usage, buffer caches, swap usage. A swap spike during stress means the OS is paging to disk, dramatically amplifying latency.
- Connection counts: Database, Redis, external APIs. Watch for connections in TIME_WAIT state — these consume file descriptors and take 60 seconds to release on Linux by default.
- Queue depths: Request queues, message queues, connection pool queues. A growing queue under steady load means the system cannot keep up.
- Error classification: Not just "errors increased" but what type — timeouts, connection refused, 502 Bad Gateway, 503 Service Unavailable. Each points to a different bottleneck.
Stress Testing in Practice
Pre-Test Checklist
- Confirm the test environment mirrors production (or document differences).
- Disable or mock external API calls to avoid stressing (or being rate-limited by) third-party services.
- Set up monitoring dashboards before the test starts — you cannot add metrics mid-test.
- Notify stakeholders so nobody panics over monitoring alerts from the test environment.
- Prepare a kill switch: know how to stop the test instantly if something goes wrong.
- Document the baseline metrics from your most recent load test for comparison.
Key Takeaways
- Stress testing finds your system's breaking point and failure mode. The goal is not to prevent failure but to ensure the failure is graceful: load shedding and degraded service, not cascading crashes.
- Use step-load patterns: increase by 10-20% every 5 minutes with hold periods. This creates clear data points for each load level.
- Connection pool exhaustion is the most common web application failure mode under stress. Set connection timeouts to fail fast rather than queue indefinitely.
- Cascading failures propagate overload across service boundaries. Circuit breakers isolate failing components and prevent propagation.
- Recovery testing is critical: continue monitoring for 10-15 minutes after load is removed. If the system doesn't return to baseline, you have a recovery problem that may be worse than the overload itself.
- Implement three layers of overload protection: rate limiting (cap throughput), load shedding (prioritize important requests), and backpressure (signal upstream to slow down).