The Three Pillars of Observability: Logs, Metrics, and Traces
Monitoring tells you when something breaks. Observability tells you why. The distinction matters because modern distributed systems fail in ways that no predetermined set of dashboards can anticipate. An observable system emits enough telemetry for engineers to ask arbitrary questions about internal state without deploying new instrumentation. That telemetry falls into three complementary categories: logs, metrics, and traces.
Each pillar captures a different dimension of system behavior. Logs record discrete events. Metrics aggregate numerical measurements over time. Traces map the journey of individual requests across service boundaries. Used in isolation, each pillar has blind spots. Combined with correlation strategies, they form a complete picture that turns debugging from guesswork into investigation.
Observability vs. Monitoring: A Critical Distinction
Monitoring is a subset of observability. A monitoring system watches known failure modes: CPU exceeds 90%, error rate spikes above 1%, disk fills past 80%. These checks work for predictable failures but collapse when systems exhibit emergent behavior that no one anticipated during setup.
Observability inverts the model. Instead of defining what to watch beforehand, an observable system generates rich, structured telemetry that supports ad-hoc exploration. When a customer reports that checkout takes 12 seconds but only on mobile Safari from Southeast Asia, an observable system lets you slice the data to isolate the cause without first building a dashboard for that exact scenario.
This distinction has practical consequences for tooling choices. Monitoring systems optimize for fixed queries against low-cardinality data. Observability platforms support high-cardinality, high-dimensionality queries that let you group and filter by arbitrary field combinations. The OpenTelemetry framework was designed specifically to produce telemetry suitable for observability rather than merely monitoring.
Pillar One: Structured Logs
Logs are the oldest form of telemetry. A log entry records a discrete event at a point in time: a request arrived, a query executed, an error occurred, a user authenticated. The challenge is not producing logs but producing logs that remain useful at scale.
Unstructured vs. Structured Logging
Traditional unstructured logs embed data inside human-readable strings. A line like 2026-07-10 14:23:01 ERROR Failed to process order 78432 for user jsmith: timeout after 30s is readable but machine-hostile. Extracting the order ID, user, or timeout value requires regex parsing that breaks whenever the message format changes.
Structured logging emits events as key-value pairs in a machine-parseable format, typically JSON:
Every field is independently queryable. You can search for all timeout errors, all errors for user jsmith, or all order-processing failures with timeouts exceeding 10 seconds. The trace_id and span_id fields connect this log entry to the broader request trace, enabling cross-pillar correlation.
Log Levels and Semantic Meaning
Effective log levels follow consistent semantic rules across all services:
| Level | Semantic Meaning | Example |
|---|---|---|
| FATAL | Process cannot continue, will exit | Failed to bind to port, corrupted database |
| ERROR | Operation failed, requires attention | Payment gateway returned 500, query timeout |
| WARN | Unexpected but handled condition | Retry succeeded on second attempt, cache miss fallback |
| INFO | Significant business or lifecycle event | Order placed, service started, deployment completed |
| DEBUG | Diagnostic detail for development | Cache key generated, SQL query text, routing decision |
The most common mistake is logging everything at INFO level, which makes it impossible to filter signal from noise in production. A well-calibrated system produces roughly 10 INFO entries per request, with DEBUG disabled unless actively investigating an issue. Alerting strategies should trigger on ERROR and FATAL entries, not WARN or INFO.
Log Aggregation Architecture
Individual log files on individual servers are useless for distributed systems. Log aggregation pipelines collect, parse, enrich, and index logs from every service instance into a centralized store. A typical pipeline follows this flow:
The collector layer is critical for reliability. Agents like Fluentd, Fluent Bit, or Vector run alongside each application instance, tail log streams, parse and enrich entries, and forward them to the aggregation layer. A buffer queue between collection and indexing protects the pipeline from backpressure during log spikes.
Pillar Two: Dimensional Metrics
Metrics are numerical measurements collected at regular intervals. Where logs record individual events, metrics aggregate behavior over time windows. A metric telling you that p99 latency was 450ms over the last five minutes is more actionable than sifting through millions of individual request logs to compute the same figure.
Metric Types
Most metrics systems support four fundamental types:
- Counters track cumulative totals that only increase: total requests served, total bytes transferred, total errors. Rate-of-change queries transform counters into throughput measurements.
- Gauges represent point-in-time values that can go up or down: current CPU usage, memory consumption, queue depth, active connections.
- Histograms sample observations into configurable buckets: request duration distribution, response size distribution. They enable percentile calculations (p50, p95, p99) without storing every individual value.
- Summaries calculate streaming percentiles on the client side. They provide precise quantiles but cannot be aggregated across instances, making them less suitable for horizontally-scaled services.
Dimensional Labels and Cardinality
Modern metrics systems attach dimensional labels (also called tags) to every measurement. A request duration metric might carry labels for service name, HTTP method, response status code, and endpoint path. These labels transform a single metric name into a multidimensional dataset that supports flexible queries.
The power of dimensional metrics comes with a trap: cardinality explosion. Every unique combination of label values creates a separate time series. A metric with labels for service (10 values), method (5 values), status (20 values), and endpoint (200 values) produces 200,000 time series. Adding a user_id label with millions of distinct values would multiply that into billions, overwhelming any metrics backend.
The Four Golden Signals
Google's SRE handbook identifies four signals that capture the health of any request-serving system:
- Latency — the time to serve a request. Separate successful and failed request latency, since a fast error is not a sign of health.
- Traffic — demand on the system, measured in requests per second, concurrent sessions, or transactions per minute.
- Errors — the rate of failed requests. Include both explicit errors (HTTP 5xx) and implicit ones (200 responses with wrong data, responses exceeding an SLO threshold).
- Saturation — how full the system is. CPU utilization, memory pressure, queue depth, connection pool exhaustion, and thread pool utilization all measure saturation of different resources.
These four signals form the minimum viable monitoring dashboard for every service. More specific metrics complement them, but the golden signals provide the starting point for any investigation. Teams that align their SLOs and error budgets with these signals create tight feedback loops between alerting and business impact.
Pillar Three: Distributed Traces
A trace represents the complete journey of a single request through a distributed system. When a user clicks "Place Order" and that click triggers calls to an API gateway, authentication service, inventory service, payment processor, and notification service, a trace connects all those interactions into a single causal chain.
Trace Anatomy
Every trace consists of spans. A span represents a single unit of work: an HTTP request, a database query, a cache lookup, a message queue publish. Each span records its start time, duration, status, and a set of attributes. Spans form a tree structure where child spans represent work initiated by parent spans.
The trace above reveals that the API gateway took 245ms total, with payment processing dominating the critical path at 95ms. The inventory service spent 28ms on a database query and only 4ms on cache lookup, suggesting the cache is working effectively. Without a trace, diagnosing why this particular request was slow would require correlating timestamps across six different service log streams.
Context Propagation
Traces work across service boundaries through context propagation. When Service A calls Service B, it injects the trace ID and parent span ID into the request headers. Service B extracts this context and uses it to create child spans that belong to the same trace. The W3C Trace Context standard defines two headers for this purpose:
The traceparent header carries the trace ID, parent span ID, and trace flags. The tracestate header allows vendors to propagate additional context without breaking interoperability. Distributed tracing fundamentals cover the implementation details of context propagation across different protocols including HTTP, gRPC, and message queues.
Sampling Strategies
Tracing every request in a high-throughput system is prohibitively expensive. A service handling 50,000 requests per second would generate millions of spans per minute. Sampling reduces this volume while preserving statistical significance:
- Head-based sampling decides at the entry point whether to trace a request. Simple probability sampling (e.g., trace 1% of requests) is easy to implement but misses rare interesting events.
- Tail-based sampling collects all spans temporarily and decides after the trace completes whether to keep it. This captures all errors and slow requests but requires a collection buffer that holds complete traces before the sampling decision.
- Priority sampling forces collection for specific conditions: all errors, all requests from specific users, all requests exceeding latency thresholds. This ensures that the most diagnostic traces are always captured.
Cross-Pillar Correlation
The real power of observability emerges when the three pillars connect. A metric alert fires when p99 latency exceeds 500ms. The engineer clicks through to the relevant time range and discovers that the latency increase correlates with increased error rates from the payment service. They pivot to traces filtered by the payment service during that window and find that a specific database query is taking 400ms instead of its usual 20ms. They follow the trace ID to the structured logs and discover that the query is hitting a table that just received a schema migration that dropped an index.
This investigative flow requires correlation identifiers that link telemetry across pillars:
Implementing Correlation
Three mechanisms enable cross-pillar navigation:
- Trace ID in logs: Every structured log entry includes the trace_id and span_id of the active trace context. This allows direct jumps from a trace timeline to the corresponding log entries, and from a log search result back to the full trace.
- Exemplars on metrics: Metrics systems like Prometheus support exemplars, which attach a sample trace_id to metric data points. When you see a latency spike on a histogram, the exemplar links directly to an example trace exhibiting that latency.
- Shared dimensions: Logs, metrics, and traces all carry consistent labels for service name, environment, region, and version. This shared vocabulary lets you pivot between pillars using the same filter criteria.
The OpenTelemetry specification defines a unified data model that produces all three telemetry types with consistent correlation identifiers. This eliminates the integration burden of wiring separate logging, metrics, and tracing libraries together.
Building an Observable System
Instrumentation Strategy
Effective instrumentation follows concentric layers. Start with automatic instrumentation that captures HTTP requests, database queries, and external calls without code changes. Add custom instrumentation for business-specific operations: order placement, search queries, recommendation generation. Finally, add diagnostic instrumentation for known trouble spots: cache eviction events, connection pool exhaustion, circuit breaker state transitions.
Each layer adds telemetry proportional to its diagnostic value. Automatic instrumentation covers 80% of debugging scenarios with zero effort. Custom business instrumentation covers 15% more. The remaining 5% requires targeted diagnostic instrumentation that you add and remove as specific issues arise.
Cost Management
Observability data is expensive to store and query. A medium-sized microservices architecture can easily generate terabytes of telemetry per day. Cost management requires deliberate decisions about retention, sampling, and tiering:
- Hot tier (7-14 days): Full-resolution data for active debugging. Fast queries, expensive storage.
- Warm tier (30-90 days): Downsampled metrics, sampled traces, filtered logs. Slower queries, moderate cost.
- Cold tier (1+ years): Aggregated metrics only. Capacity planning, trend analysis, compliance.
Teams building their first observability stack often underestimate storage costs. A service-level monitoring approach helps by focusing instrumentation on the services and interactions that matter most, rather than instrumenting everything uniformly.
Common Anti-Patterns
Several patterns undermine observability despite appearing productive:
- Dashboard-driven observability: Building dashboards before understanding failure modes produces walls of green charts that turn red simultaneously during incidents, providing no diagnostic signal.
- Log-only debugging: Relying exclusively on logs for distributed systems forces engineers to manually correlate timestamps across services, a process that scales poorly and misses ordering nuances.
- Metrics without context: Metrics that aggregate away all dimensionality (e.g., total error count with no service or endpoint labels) answer none of the follow-up questions that alerts generate.
- Trace sampling too aggressively: Sampling at 0.01% captures only happy-path requests and misses the errors and edge cases that engineers actually need to debug.
- Siloed pillars: Running separate, unlinked systems for logs, metrics, and traces defeats the investigative workflow that makes observability valuable. Without correlation, engineers context-switch between tools and lose the thread.
Avoiding these anti-patterns requires treating observability as a first-class engineering concern rather than an afterthought bolted onto existing systems. Teams that invest in structured, correlated telemetry from the beginning spend dramatically less time on incident response and maintain faster regression detection cycles.
Key Takeaways
- Observability differs from monitoring in its support for ad-hoc, high-cardinality queries against system telemetry.
- Logs record discrete events and support high-cardinality investigations. Structure them as key-value pairs with consistent schemas.
- Metrics aggregate numerical measurements over time. Keep label cardinality bounded and align with the four golden signals.
- Traces map request journeys across services. Use sampling strategies appropriate to traffic volume while preserving error and latency outliers.
- Cross-pillar correlation through trace IDs, exemplars, and shared dimensions transforms three data streams into a unified investigative tool.