Home › Observability & SRE › Chaos Engineering for Performance

Chaos Engineering for Performance: Breaking Systems to Build Resilience

Your system handles 10,000 requests per second at steady state. Then a cache node fails. Without the cache, database query latency jumps from 2ms to 200ms. Connection pools exhaust. Timeouts cascade through dependent services. Within three minutes, your entire platform is down. You discover this at 2 AM during Black Friday.

Chaos engineering prevents this scenario by deliberately injecting failures during controlled conditions. Rather than waiting for production incidents to reveal weaknesses, you proactively test how your system behaves under adverse conditions—and fix the problems before they find you.

Principles of Chaos Engineering

Chaos engineering is not random destruction. It follows the scientific method applied to distributed systems. Netflix, which pioneered the discipline, defines it through four principles:

  1. Define steady state: Establish measurable indicators of normal system behavior. Request rate, error rate, latency percentiles, and throughput collectively define what "healthy" looks like. These become your SLIs.
  2. Hypothesize about steady state: Form a prediction about how the system will behave under a specific perturbation. "If one cache node fails, latency will increase by no more than 50ms and error rate will stay below 0.1%."
  3. Introduce real-world events: Inject the failure: kill the cache node, add network latency, fill a disk, terminate a process. Use realistic failures, not theoretical edge cases.
  4. Observe the difference: Compare actual behavior against the hypothesis. Did the system gracefully degrade, or did it cascade? Was the hypothesis confirmed or disproved?
Chaos Experiment Workflow 1. Steady State Define normal behavior metrics 2. Hypothesize Predict impact of failure 3. Inject Fault Execute the experiment 4. Observe Compare actual vs predicted 5. Fix Remediate weaknesses Repeat with new hypothesis Hypothesis Confirmed System resilient. Try harder. Hypothesis Disproved Weakness found. Fix and re-test.

Performance-Specific Fault Injection

Traditional chaos engineering focuses on availability: killing instances, partitioning networks, corrupting data. Performance chaos engineering targets degradation scenarios that do not cause outright failure but erode user experience:

Latency Injection

Add artificial latency to network calls between services. This tests timeout configurations, circuit breaker thresholds, and user-facing degradation handling. A 500ms delay added to database queries reveals whether your service correctly returns partial results, shows loading states, or simply times out with a generic error.

# Toxiproxy latency injection toxiproxy-cli toxic add \ --type latency \ --attribute latency=500 \ --attribute jitter=200 \ --toxicName db-slowdown \ database-proxy # Litmus Chaos - network latency experiment apiVersion: litmuschaos.io/v1alpha1 kind: ChaosEngine metadata: name: latency-experiment spec: appinfo: appns: production applabel: app=order-service experiments: - name: pod-network-latency spec: components: env: - name: NETWORK_LATENCY value: '500' - name: JITTER value: '200' - name: DESTINATION_IPS value: '10.0.1.50' # database IP

CPU and Memory Pressure

Simulate resource contention by consuming CPU cycles or allocating memory on target pods. This tests autoscaling responsiveness, garbage collection behavior under memory pressure, and the effectiveness of resource limits. A service that handles 500 RPS at normal CPU may drop to 50 RPS when a noisy neighbor consumes 70% of available CPU.

Network Degradation

Inject packet loss, bandwidth throttling, and DNS resolution delays. These simulate real network conditions: lossy WiFi connections, congested cross-region links, CDN failures. Network degradation reveals assumptions baked into retry logic and timeout configurations across your service mesh.

Dependency Degradation

Slow down or partially fail an upstream dependency. If your service calls a recommendation engine, inject a 50% error rate on that call and verify that the primary flow still works—perhaps without recommendations. This validates graceful degradation patterns and fallback logic.

GameDay Exercises

GameDays are structured events where teams deliberately break their own systems to test resilience. Unlike automated chaos experiments, GameDays involve human operators, cross-team coordination, and real-time decision making.

Planning a GameDay

  1. Select a scenario: Choose a realistic failure that the team believes the system can handle. The goal is validation, not surprise. "Cache cluster failure during peak traffic" is a good first scenario.
  2. Define success criteria: What does graceful degradation look like? P99 latency stays under 500ms, error rate stays below 1%, no data loss, recovery within 5 minutes.
  3. Prepare the abort procedure: Define clear criteria for halting the exercise. If error rate exceeds 5% or any data integrity issue is detected, immediately roll back the fault injection.
  4. Assign roles: Experiment driver (injects faults), observer (monitors dashboards), facilitator (coordinates communication), and safety officer (triggers abort if needed).
  5. Notify stakeholders: Inform customer support, product management, and any team with dependencies. Even controlled experiments can produce customer-visible effects.
Blast radius control: Start with non-production environments. Graduate to production only with canary-style experiments that affect a small percentage of traffic. Never inject failures that could cause data loss or corruption without verified backup and recovery procedures.

Running the Exercise

During the GameDay, the experiment driver injects faults according to the plan while the team monitors impact on service-level indicators. Record everything: what was injected, when, what the dashboards showed, what alerts fired, what actions operators took. This record becomes the basis for the post-exercise review.

Chaos in CI/CD

Integrate chaos experiments into your deployment pipeline. Every release candidate runs through a suite of fault injection tests before reaching production. This catches resilience regressions the same way unit tests catch functional regressions.

Pipeline StageChaos TestPass Criteria
IntegrationDependency timeout (500ms added)P99 < 2s, error rate < 0.5%
IntegrationDependency 10% error rateFallback activated, primary flow works
StagingPod kill (random 1 of N)Zero downtime, auto-recovery < 30s
StagingCPU stress (80% for 2 min)RPS drop < 20%, no OOM kills
CanaryNetwork partition (1 AZ)Cross-AZ failover < 10s

Measuring Resilience

Track resilience improvements quantitatively over time:

Common Performance Chaos Scenarios

Cache Failure

Disable or slow down cache layers to test the system's behavior under cache miss storms. Verify that database connection pools can handle the increased load, that rate limiters prevent thundering herds, and that cache warming procedures execute correctly.

Downstream Service Slowdown

Add latency to a critical downstream service. Verify that circuit breakers trip within expected timeframes, that fallback responses are served, and that the calling service does not exhaust its thread pool or connection pool waiting for slow responses.

Resource Exhaustion

Consume available memory, disk space, or file descriptors on a target host. Verify that resource limits prevent cascade, that eviction policies work correctly, and that monitoring alerts fire before resource exhaustion causes an outage.

Clock Skew

Advance or retard the system clock on a subset of nodes. This reveals assumptions about timestamp ordering in distributed systems: token expiration, cache TTLs, log correlation, and incident timeline reconstruction.

Getting Started

Start small. Run your first chaos experiment in a staging environment with a single, reversible fault injection. As confidence grows, graduate to production experiments with increasing scope:

  1. Week 1: Inventory critical dependencies and their expected failure modes
  2. Week 2: Run one latency injection experiment in staging
  3. Week 3: Add chaos tests to CI/CD for the most critical service
  4. Month 2: Run first GameDay with a single team
  5. Month 3: Introduce production chaos experiments with canary scope (1% of traffic)
  6. Quarter 2: Cross-team GameDay exercises, automated chaos suite in CI/CD

Key Takeaways