Distributed Tracing Fundamentals: OpenTelemetry, Spans, and Context Propagation
When a user clicks "Place Order" in a modern web application, their request may traverse an API gateway, authentication service, inventory service, payment processor, notification system, and fulfillment engine — each running as an independent microservice, potentially in different cloud regions. If that request takes 8 seconds instead of the expected 800 milliseconds, which service is responsible?
Distributed tracing answers this question by recording the complete journey of a request across every service it touches. Each operation contributes a timestamped record called a "span," and spans are linked through a shared identifier — the trace ID — forming a complete picture of what happened, where it happened, and how long each piece took.
This guide covers the core concepts of distributed tracing, the OpenTelemetry standard, context propagation mechanisms, span modeling best practices, and sampling strategies that balance observability with cost. For the broader monitoring context, see our APM fundamentals article.
Trace Anatomy: Traces, Spans, and Context
A trace is a tree of spans representing the complete lifecycle of a single request. Every trace has exactly one root span — the initial operation that started the request — and zero or more child spans representing downstream operations.
Span Attributes
Each span carries structured metadata that makes it queryable and analyzable:
- Trace ID: A globally unique 128-bit identifier shared by all spans in the trace
- Span ID: A 64-bit identifier unique to this span within the trace
- Parent Span ID: Links this span to its parent, forming the tree structure
- Operation name: What the span represents (e.g.,
HTTP POST /orders,db.query) - Start and end timestamps: When the operation began and completed
- Status: OK, ERROR, or UNSET
- Attributes: Key-value pairs (e.g.,
http.method=POST,db.statement=SELECT...) - Events: Timestamped annotations within the span's lifetime (e.g., exception events)
Context Propagation
The mechanism that links spans across service boundaries is context propagation. When Service A calls Service B, it injects the current trace context (trace ID, parent span ID, trace flags) into the outgoing request — typically as HTTP headers. Service B extracts this context and uses it as the parent for its own spans.
W3C Trace Context
The W3C Trace Context specification defines two standardized HTTP headers:
# traceparent header format:
# version-traceId-parentId-traceFlags
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
# tracestate header (vendor-specific key-value pairs):
tracestate: congo=t61rcWkgMzE,rojo=00f067aa0ba902b7
The traceparent header carries the essential propagation data: version (00), trace ID (32 hex characters), parent span ID (16 hex characters), and trace flags (where 01 means "sampled"). The tracestate header allows vendors to attach additional context without breaking interoperability.
Propagation in Practice
OpenTelemetry auto-instrumentation handles context propagation automatically for supported frameworks. When you instrument an HTTP client library (like requests in Python or axios in Node.js), the instrumentation injects traceparent headers on outgoing requests. Server-side instrumentation extracts these headers and creates child spans.
// Node.js: Manual context propagation with OpenTelemetry
const { context, propagation, trace } = require('@opentelemetry/api');
// Injecting context into outgoing request
const headers = {};
propagation.inject(context.active(), headers);
// headers now contains { traceparent: '00-abc...-def...-01' }
// Extracting context from incoming request
const extractedContext = propagation.extract(
context.active(),
incomingRequest.headers
);
context.with(extractedContext, () => {
const span = tracer.startSpan('process-order');
// This span is now a child of the extracted context
});
OpenTelemetry Architecture
OpenTelemetry provides the standard instrumentation layer for distributed tracing. Its architecture consists of three key components:
- SDK: Language-specific libraries that create and manage spans. The SDK handles span lifecycle, attribute management, and export. Each language has its own SDK (Python, Java, Go, JavaScript, .NET, Ruby, etc.).
- Exporters: Plugins that send span data to a backend. OTLP (OpenTelemetry Protocol) is the native export format, but exporters exist for Jaeger, Zipkin, and vendor-specific formats.
- Collector: A standalone process that receives, processes, and re-exports telemetry data. The Collector decouples instrumentation from the backend, allowing you to change backends without modifying application code.
# OpenTelemetry Collector configuration (otel-collector.yaml)
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 5s
send_batch_size: 1024
tail_sampling:
decision_wait: 10s
policies:
- name: errors
type: status_code
status_code: {status_codes: [ERROR]}
- name: slow-traces
type: latency
latency: {threshold_ms: 2000}
- name: probabilistic
type: probabilistic
probabilistic: {sampling_percentage: 10}
exporters:
otlp/jaeger:
endpoint: jaeger:4317
tls:
insecure: true
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch, tail_sampling]
exporters: [otlp/jaeger]
Sampling Strategies
In production systems handling millions of requests per minute, capturing every trace is neither practical nor affordable. Sampling determines which traces are recorded and which are dropped.
| Strategy | How It Works | Pros | Cons |
|---|---|---|---|
| Head-based probabilistic | Decision made at trace start. Each trace has X% chance of being sampled. | Simple, consistent, predictable cost | May miss rare errors and outliers |
| Head-based rate limiting | Sample N traces per second, drop the rest | Predictable volume regardless of traffic | Under-represents bursts |
| Tail-based | Decision made after trace completes. Keep errors, slow traces, drop normal ones. | Captures all interesting traces | Requires buffering complete traces; higher memory cost |
| Adaptive / dynamic | Adjusts sampling rate based on traffic volume or error rate | Balances cost and coverage automatically | Complex to configure; can miss initial errors in a burst |
Tail-based sampling is the gold standard for production systems. By waiting until a trace completes, the sampler can make informed decisions: keep all traces with errors, keep traces with latency above the p99 threshold, and probabilistically sample the rest. This ensures you never miss a diagnostic trace while controlling storage costs. The OpenTelemetry Collector's tail_sampling processor implements this pattern.
Span Modeling Best Practices
How you structure spans determines how useful your traces are for debugging. Poor span modeling produces traces that are either too noisy (hundreds of trivial spans) or too sparse (missing the operation that matters).
What Deserves a Span
- Service boundaries: Every incoming and outgoing RPC, HTTP call, or message queue interaction
- Database operations: Queries, transactions, connection acquisition
- External API calls: Third-party services like payment processors or email providers
- Significant business operations: Order validation, fraud check, inventory reservation — operations that a human would want to see in the trace when debugging
What Does Not Need a Span
- In-memory computations: String parsing, data transformation, business logic that takes microseconds
- Loop iterations: Avoid creating a span per item in a loop — aggregate instead
- Framework internals: Middleware chains, filter pipelines — unless they're performance-relevant
# Python: Well-structured custom spans
from opentelemetry import trace
tracer = trace.get_tracer("order-service")
def process_order(order):
with tracer.start_as_current_span("validate-order") as span:
span.set_attribute("order.id", order.id)
span.set_attribute("order.item_count", len(order.items))
validate(order)
with tracer.start_as_current_span("check-inventory") as span:
span.set_attribute("order.id", order.id)
available = inventory_client.check(order.items)
span.set_attribute("inventory.all_available", available)
with tracer.start_as_current_span("process-payment") as span:
span.set_attribute("order.id", order.id)
span.set_attribute("payment.method", order.payment_method)
result = payment_client.charge(order)
span.set_attribute("payment.status", result.status)
Correlating Traces with Metrics and Logs
Traces are most powerful when correlated with the other observability pillars described in our APM overview. The trace ID serves as the correlation key:
- Traces → Logs: Include the trace ID in every log entry. When investigating a failed trace, search logs by trace ID to find the exact error message, stack trace, and surrounding context.
- Metrics → Traces: Use OpenTelemetry "exemplars" — sample trace IDs attached to metric data points. When a metric shows a latency spike, click through to an exemplar trace that demonstrates the cause.
- Frontend → Backend: Core Web Vitals measurements can be linked to backend traces when the browser includes a trace ID in its requests, connecting a poor LCP score to the specific backend bottleneck.
Common Distributed Tracing Patterns
Fan-Out/Fan-In
A service calls multiple downstream services in parallel (fan-out), then aggregates their responses (fan-in). In the trace, this appears as multiple sibling child spans under a single parent. The parent span's duration includes the longest child, making it easy to identify which parallel call is the bottleneck.
Async Message Processing
For message queue-based architectures (Kafka, RabbitMQ, SQS), trace context is injected into message headers by the producer and extracted by the consumer. This creates a "link" between the producer's trace and the consumer's trace, even though the message may be processed minutes or hours later.
Retry and Circuit Breaker
When a service retries a failed request or a circuit breaker opens, each retry attempt should be a new child span with attributes indicating the retry count. This makes retry storms visible in traces and helps identify services that are causing cascading failures.
Debugging with Traces: A Workflow
- Symptom detection: A metric alert fires — p95 latency for
/checkoutexceeds 3 seconds - Find exemplar traces: From the metric dashboard, click an exemplar to open a trace that exhibits the high latency
- Identify the bottleneck: In the waterfall view, locate the span with the longest duration — the widest bar in the chart
- Examine attributes: The bottleneck span's attributes reveal the cause — a slow database query, a third-party API timeout, or resource contention
- Correlate with logs: Search logs by the trace ID for detailed error messages, stack traces, or query plans
- Verify the fix: After deploying a fix, compare trace durations before and after using performance measurement tools
Key Takeaways
- A trace is a tree of spans representing a request's journey across services
- W3C Trace Context (
traceparentheader) is the interoperability standard for context propagation - OpenTelemetry provides vendor-neutral instrumentation across all major languages
- Tail-based sampling captures all error and latency-outlier traces while controlling cost
- Create spans for service boundaries, database operations, and significant business logic — not for trivial in-memory operations
- Correlate traces with logs (via trace ID) and metrics (via exemplars) for full observability