Service-Level Monitoring: From Health Checks to User Journeys
A service that responds with HTTP 200 is not necessarily healthy. The database connection pool may be exhausted. The cache may be serving stale data. An upstream dependency may be degraded, causing silent data corruption. Simple health checks verify that a process is running, but they say nothing about whether the service is actually fulfilling its purpose for users.
Effective service monitoring operates across multiple layers, from basic liveness probes up through business-level user journey verification. Each layer catches a different class of failure, and a mature monitoring stack needs all of them working together.
The Monitoring Pyramid
Layer 1: Health Checks
Health checks are the foundation. They answer two questions: is the process alive (liveness), and is it ready to serve traffic (readiness)?
Liveness Probes
A liveness probe verifies that the process is running and responsive. In Kubernetes, a failing liveness probe triggers a container restart. Keep liveness probes simple—they should verify that the application event loop is responsive, not that every dependency is available. An overly thorough liveness check that queries the database turns a database outage into a cascading restart storm.
Readiness Probes
A readiness probe determines whether a pod should receive traffic. Unlike liveness, readiness probes should check critical dependencies: database connectivity, cache availability, configuration loaded. A pod that fails readiness is removed from the service load balancer but not restarted, allowing it to recover when the dependency returns.
Layer 2: Infrastructure Monitoring
Infrastructure metrics capture resource utilization and saturation. The USE method (Utilization, Saturation, Errors) provides a systematic framework for infrastructure monitoring:
| Resource | Utilization | Saturation | Errors |
|---|---|---|---|
| CPU | % time busy | Run queue length | Machine check exceptions |
| Memory | % used | Swap usage, OOM events | Allocation failures |
| Disk I/O | % device busy | I/O queue depth | Read/write errors |
| Network | Bandwidth usage | TCP retransmits | Interface errors, drops |
| Connection Pool | Active/total ratio | Wait queue length | Timeout errors |
Node Exporter (Linux), Windows Exporter, and cAdvisor (containers) expose these metrics for Prometheus scraping. Track cardinality carefully when adding infrastructure labels.
Layer 3: API and Endpoint Monitoring
API monitoring verifies that endpoints respond correctly under real conditions. This goes beyond health checks to validate response content, measure latency distributions, and track error rates per endpoint.
The RED method (Rate, Errors, Duration) provides the essential metrics for every service endpoint:
- Rate: Requests per second. Sudden drops indicate upstream issues or routing changes. Unexpected spikes suggest traffic events or retry storms.
- Errors: Error rate as a percentage of total requests. Distinguish between client errors (4xx, which may be normal) and server errors (5xx, which always warrant investigation).
- Duration: Response time distribution, not just averages. P50 shows typical experience, P95 shows slow-tail impact, P99 reveals worst-case scenarios affecting 1% of users.
Response Validation
Status code monitoring catches outright failures but misses subtle degradation. Response validation checks verify that the response body contains expected content. An API might return 200 OK with an empty data array when the database query silently fails. Validate response schema, critical field presence, and data freshness.
Layer 4: Synthetic Monitoring
Synthetic monitoring uses automated scripts that simulate user interactions from external vantage points. Unlike real-user monitoring, synthetic tests run continuously regardless of traffic volume, providing consistent baseline measurements and detecting issues during low-traffic periods.
Single-Step Checks
Single-step synthetics test individual endpoints from multiple geographic locations. They measure DNS resolution time, TCP connection time, TLS handshake duration, time to first byte, and total response time. Run checks from every region where users are located to detect regional infrastructure issues.
Multi-Step Transactions
Multi-step synthetics execute scripted user flows: login, search for a product, add to cart, proceed to checkout. Each step validates both the response and the state transitions. If step 3 (add to cart) fails, the monitor reports which specific interaction broke without masking the root cause behind a generic "flow failed" alert.
Layer 5: User Journey Monitoring
User journey monitoring tracks complete business workflows as experienced by real users. While synthetic monitoring simulates users, journey monitoring tracks actual user sessions to measure conversion funnels, error rates at each step, and SLO compliance from the user's perspective.
Journey Definition
Define journeys around business-critical paths. For an e-commerce platform, the primary journey is: homepage → search → product detail → add to cart → checkout → order confirmation. For a SaaS application: login → dashboard load → key feature interaction → data save.
Each journey step has measurable success criteria:
- The step completed (no error, expected state transition)
- The step completed within the performance budget
- The step produced the correct output (validation passed)
Funnel Analysis
Track drop-off rates between journey steps. If 10,000 users start checkout but only 7,000 reach payment, the 30% drop-off deserves investigation. Correlate drop-offs with performance metrics: do users who experience > 3s load times abandon at 5x the rate of users under 1s? This connects performance data directly to business outcomes.
Alerting Architecture
Alert Hierarchy
Not every monitoring signal justifies waking someone up. Structure alerts in tiers:
| Tier | Signal Source | Response | Example |
|---|---|---|---|
| P1 - Page | User journey failures, SLO burn rate | Immediate oncall page | Checkout flow 0% success rate |
| P2 - Urgent | Synthetic failures, high error rates | Slack alert, investigate within 30 min | API P99 latency 5x above baseline |
| P3 - Warning | Infrastructure saturation, health degradation | Ticket, fix within business hours | Connection pool at 80% utilization |
| P4 - Info | Capacity trends, metric anomalies | Dashboard review, planning | Disk usage growth trend |
Reducing Alert Fatigue
Alert fatigue—the state where responders ignore alerts because most are false positives—is more dangerous than missing instrumentation. Teams that receive more than 5 actionable alerts per oncall shift start to lose response quality. Reduce noise by:
- Alerting on symptoms (user impact), not causes (CPU spike). High CPU that does not affect latency or error rate does not need a page.
- Using multi-signal confirmation. A latency increase confirmed by both synthetic tests and real-user metrics is more credible than either alone.
- Setting appropriate durations. A 30-second CPU spike is not an alert. Five minutes of sustained error rate increase is.
- Reviewing alert history monthly. Delete alerts that never fired and investigate alerts that fired but required no action.
Uptime Calculation
Uptime percentages are meaningless without precise definitions. Define what constitutes "up" at each monitoring layer:
- Infrastructure uptime: The service process is running and health checks pass. This is the easiest to achieve and the least meaningful. 99.99% infrastructure uptime is table stakes.
- API uptime: The service responds with correct data within latency SLOs. More meaningful but still incomplete.
- User journey uptime: Complete business flows succeed from the user's perspective. This is the metric that matters. A service with 99.99% API uptime but 95% checkout success rate is failing its users.
Key Takeaways
- Effective monitoring operates across five layers: health checks, infrastructure, API endpoints, synthetic tests, and user journeys. Each catches a different failure class.
- Separate liveness probes (restart on failure) from readiness probes (remove from load balancer). Keep liveness simple to avoid restart storms.
- Use the RED method (Rate, Errors, Duration) for every service endpoint. Track latency distributions, not just averages.
- Synthetic monitoring detects issues during low-traffic periods and from the user's geographic perspective. Run multi-step transaction scripts for critical flows.
- Structure alerts in tiers based on user impact. Alert on symptoms, use multi-signal confirmation, and review alert effectiveness monthly.