Home › Observability & SRE › Service-Level Monitoring

Service-Level Monitoring: From Health Checks to User Journeys

A service that responds with HTTP 200 is not necessarily healthy. The database connection pool may be exhausted. The cache may be serving stale data. An upstream dependency may be degraded, causing silent data corruption. Simple health checks verify that a process is running, but they say nothing about whether the service is actually fulfilling its purpose for users.

Effective service monitoring operates across multiple layers, from basic liveness probes up through business-level user journey verification. Each layer catches a different class of failure, and a mature monitoring stack needs all of them working together.

The Monitoring Pyramid

Multi-Layer Monitoring Architecture User Journeys End-to-end business flows Synthetic Monitoring Proactive multi-step checks API & Endpoint Monitoring HTTP status, latency, response validation Infrastructure Monitoring CPU, memory, disk, network, process count Health Checks (Liveness & Readiness) Highest value Hardest to build Foundation Easiest to implement

Layer 1: Health Checks

Health checks are the foundation. They answer two questions: is the process alive (liveness), and is it ready to serve traffic (readiness)?

Liveness Probes

A liveness probe verifies that the process is running and responsive. In Kubernetes, a failing liveness probe triggers a container restart. Keep liveness probes simple—they should verify that the application event loop is responsive, not that every dependency is available. An overly thorough liveness check that queries the database turns a database outage into a cascading restart storm.

# Kubernetes liveness probe - simple HTTP check livenessProbe: httpGet: path: /healthz port: 8080 initialDelaySeconds: 10 periodSeconds: 15 timeoutSeconds: 3 failureThreshold: 3

Readiness Probes

A readiness probe determines whether a pod should receive traffic. Unlike liveness, readiness probes should check critical dependencies: database connectivity, cache availability, configuration loaded. A pod that fails readiness is removed from the service load balancer but not restarted, allowing it to recover when the dependency returns.

# Readiness probe - checks dependencies readinessProbe: httpGet: path: /ready port: 8080 initialDelaySeconds: 5 periodSeconds: 10 timeoutSeconds: 5 failureThreshold: 2 # /ready endpoint implementation @app.get("/ready") async def readiness(): checks = { "database": await check_db_connection(), "cache": await check_redis_connection(), "config": config_loaded, } all_healthy = all(checks.values()) status_code = 200 if all_healthy else 503 return JSONResponse(checks, status_code=status_code)
Key distinction: Liveness probes restart the container on failure. Readiness probes remove it from load balancing. Use liveness for "is the process stuck?" and readiness for "can it serve users right now?"

Layer 2: Infrastructure Monitoring

Infrastructure metrics capture resource utilization and saturation. The USE method (Utilization, Saturation, Errors) provides a systematic framework for infrastructure monitoring:

ResourceUtilizationSaturationErrors
CPU% time busyRun queue lengthMachine check exceptions
Memory% usedSwap usage, OOM eventsAllocation failures
Disk I/O% device busyI/O queue depthRead/write errors
NetworkBandwidth usageTCP retransmitsInterface errors, drops
Connection PoolActive/total ratioWait queue lengthTimeout errors

Node Exporter (Linux), Windows Exporter, and cAdvisor (containers) expose these metrics for Prometheus scraping. Track cardinality carefully when adding infrastructure labels.

Layer 3: API and Endpoint Monitoring

API monitoring verifies that endpoints respond correctly under real conditions. This goes beyond health checks to validate response content, measure latency distributions, and track error rates per endpoint.

The RED method (Rate, Errors, Duration) provides the essential metrics for every service endpoint:

Response Validation

Status code monitoring catches outright failures but misses subtle degradation. Response validation checks verify that the response body contains expected content. An API might return 200 OK with an empty data array when the database query silently fails. Validate response schema, critical field presence, and data freshness.

Layer 4: Synthetic Monitoring

Synthetic monitoring uses automated scripts that simulate user interactions from external vantage points. Unlike real-user monitoring, synthetic tests run continuously regardless of traffic volume, providing consistent baseline measurements and detecting issues during low-traffic periods.

Single-Step Checks

Single-step synthetics test individual endpoints from multiple geographic locations. They measure DNS resolution time, TCP connection time, TLS handshake duration, time to first byte, and total response time. Run checks from every region where users are located to detect regional infrastructure issues.

Multi-Step Transactions

Multi-step synthetics execute scripted user flows: login, search for a product, add to cart, proceed to checkout. Each step validates both the response and the state transitions. If step 3 (add to cart) fails, the monitor reports which specific interaction broke without masking the root cause behind a generic "flow failed" alert.

# Synthetic monitoring script (Playwright-based) async def checkout_flow(page): # Step 1: Load homepage await page.goto('https://store.example.com') assert await page.title() == 'Example Store' # Step 2: Search for product await page.fill('[data-testid="search"]', 'wireless headphones') await page.click('[data-testid="search-submit"]') results = await page.locator('.product-card').count() assert results > 0, "No search results returned" # Step 3: Add to cart await page.click('.product-card:first-child .add-to-cart') cart_count = await page.text_content('.cart-badge') assert cart_count == '1', f"Cart shows {cart_count}, expected 1" # Step 4: Begin checkout await page.click('[data-testid="checkout"]') assert '/checkout' in page.url

Layer 5: User Journey Monitoring

User journey monitoring tracks complete business workflows as experienced by real users. While synthetic monitoring simulates users, journey monitoring tracks actual user sessions to measure conversion funnels, error rates at each step, and SLO compliance from the user's perspective.

Journey Definition

Define journeys around business-critical paths. For an e-commerce platform, the primary journey is: homepage → search → product detail → add to cart → checkout → order confirmation. For a SaaS application: login → dashboard load → key feature interaction → data save.

Each journey step has measurable success criteria:

Funnel Analysis

Track drop-off rates between journey steps. If 10,000 users start checkout but only 7,000 reach payment, the 30% drop-off deserves investigation. Correlate drop-offs with performance metrics: do users who experience > 3s load times abandon at 5x the rate of users under 1s? This connects performance data directly to business outcomes.

Alerting Architecture

Alert Hierarchy

Not every monitoring signal justifies waking someone up. Structure alerts in tiers:

TierSignal SourceResponseExample
P1 - PageUser journey failures, SLO burn rateImmediate oncall pageCheckout flow 0% success rate
P2 - UrgentSynthetic failures, high error ratesSlack alert, investigate within 30 minAPI P99 latency 5x above baseline
P3 - WarningInfrastructure saturation, health degradationTicket, fix within business hoursConnection pool at 80% utilization
P4 - InfoCapacity trends, metric anomaliesDashboard review, planningDisk usage growth trend

Reducing Alert Fatigue

Alert fatigue—the state where responders ignore alerts because most are false positives—is more dangerous than missing instrumentation. Teams that receive more than 5 actionable alerts per oncall shift start to lose response quality. Reduce noise by:

Uptime Calculation

Uptime percentages are meaningless without precise definitions. Define what constitutes "up" at each monitoring layer:

Key Takeaways