Home›APM›Distributed Tracing

Distributed Tracing Fundamentals: OpenTelemetry, Spans, and Context Propagation

Elena KowalskiSeptember 16, 202614 min read

When a user clicks "Place Order" in a modern web application, their request may traverse an API gateway, authentication service, inventory service, payment processor, notification system, and fulfillment engine — each running as an independent microservice, potentially in different cloud regions. If that request takes 8 seconds instead of the expected 800 milliseconds, which service is responsible?

Distributed tracing answers this question by recording the complete journey of a request across every service it touches. Each operation contributes a timestamped record called a "span," and spans are linked through a shared identifier — the trace ID — forming a complete picture of what happened, where it happened, and how long each piece took.

This guide covers the core concepts of distributed tracing, the OpenTelemetry standard, context propagation mechanisms, span modeling best practices, and sampling strategies that balance observability with cost. For the broader monitoring context, see our APM fundamentals article.

Trace Anatomy: Traces, Spans, and Context

A trace is a tree of spans representing the complete lifecycle of a single request. Every trace has exactly one root span — the initial operation that started the request — and zero or more child spans representing downstream operations.

Trace ID: abc-123-def 0ms 500ms 1000ms API Gateway — POST /orders (980ms) Auth (95ms) Order Service (620ms) Inventory (210ms) Payment (340ms) DB (45ms) Stripe API (280ms) Notify (85ms) Waterfall View of a Distributed Trace Each bar is a span. Nesting shows parent-child relationships. Width represents duration. The Payment → Stripe span is the bottleneck.

Span Attributes

Each span carries structured metadata that makes it queryable and analyzable:

Context Propagation

The mechanism that links spans across service boundaries is context propagation. When Service A calls Service B, it injects the current trace context (trace ID, parent span ID, trace flags) into the outgoing request — typically as HTTP headers. Service B extracts this context and uses it as the parent for its own spans.

W3C Trace Context

The W3C Trace Context specification defines two standardized HTTP headers:

# traceparent header format:
# version-traceId-parentId-traceFlags
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01

# tracestate header (vendor-specific key-value pairs):
tracestate: congo=t61rcWkgMzE,rojo=00f067aa0ba902b7

The traceparent header carries the essential propagation data: version (00), trace ID (32 hex characters), parent span ID (16 hex characters), and trace flags (where 01 means "sampled"). The tracestate header allows vendors to attach additional context without breaking interoperability.

Propagation in Practice

OpenTelemetry auto-instrumentation handles context propagation automatically for supported frameworks. When you instrument an HTTP client library (like requests in Python or axios in Node.js), the instrumentation injects traceparent headers on outgoing requests. Server-side instrumentation extracts these headers and creates child spans.

// Node.js: Manual context propagation with OpenTelemetry
const { context, propagation, trace } = require('@opentelemetry/api');

// Injecting context into outgoing request
const headers = {};
propagation.inject(context.active(), headers);
// headers now contains { traceparent: '00-abc...-def...-01' }

// Extracting context from incoming request
const extractedContext = propagation.extract(
  context.active(),
  incomingRequest.headers
);
context.with(extractedContext, () => {
  const span = tracer.startSpan('process-order');
  // This span is now a child of the extracted context
});

OpenTelemetry Architecture

OpenTelemetry provides the standard instrumentation layer for distributed tracing. Its architecture consists of three key components:

  1. SDK: Language-specific libraries that create and manage spans. The SDK handles span lifecycle, attribute management, and export. Each language has its own SDK (Python, Java, Go, JavaScript, .NET, Ruby, etc.).
  2. Exporters: Plugins that send span data to a backend. OTLP (OpenTelemetry Protocol) is the native export format, but exporters exist for Jaeger, Zipkin, and vendor-specific formats.
  3. Collector: A standalone process that receives, processes, and re-exports telemetry data. The Collector decouples instrumentation from the backend, allowing you to change backends without modifying application code.
# OpenTelemetry Collector configuration (otel-collector.yaml)
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 5s
    send_batch_size: 1024
  tail_sampling:
    decision_wait: 10s
    policies:
      - name: errors
        type: status_code
        status_code: {status_codes: [ERROR]}
      - name: slow-traces
        type: latency
        latency: {threshold_ms: 2000}
      - name: probabilistic
        type: probabilistic
        probabilistic: {sampling_percentage: 10}

exporters:
  otlp/jaeger:
    endpoint: jaeger:4317
    tls:
      insecure: true

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch, tail_sampling]
      exporters: [otlp/jaeger]

Sampling Strategies

In production systems handling millions of requests per minute, capturing every trace is neither practical nor affordable. Sampling determines which traces are recorded and which are dropped.

StrategyHow It WorksProsCons
Head-based probabilisticDecision made at trace start. Each trace has X% chance of being sampled.Simple, consistent, predictable costMay miss rare errors and outliers
Head-based rate limitingSample N traces per second, drop the restPredictable volume regardless of trafficUnder-represents bursts
Tail-basedDecision made after trace completes. Keep errors, slow traces, drop normal ones.Captures all interesting tracesRequires buffering complete traces; higher memory cost
Adaptive / dynamicAdjusts sampling rate based on traffic volume or error rateBalances cost and coverage automaticallyComplex to configure; can miss initial errors in a burst

Tail-based sampling is the gold standard for production systems. By waiting until a trace completes, the sampler can make informed decisions: keep all traces with errors, keep traces with latency above the p99 threshold, and probabilistically sample the rest. This ensures you never miss a diagnostic trace while controlling storage costs. The OpenTelemetry Collector's tail_sampling processor implements this pattern.

Span Modeling Best Practices

How you structure spans determines how useful your traces are for debugging. Poor span modeling produces traces that are either too noisy (hundreds of trivial spans) or too sparse (missing the operation that matters).

What Deserves a Span

What Does Not Need a Span

# Python: Well-structured custom spans
from opentelemetry import trace

tracer = trace.get_tracer("order-service")

def process_order(order):
    with tracer.start_as_current_span("validate-order") as span:
        span.set_attribute("order.id", order.id)
        span.set_attribute("order.item_count", len(order.items))
        validate(order)

    with tracer.start_as_current_span("check-inventory") as span:
        span.set_attribute("order.id", order.id)
        available = inventory_client.check(order.items)
        span.set_attribute("inventory.all_available", available)

    with tracer.start_as_current_span("process-payment") as span:
        span.set_attribute("order.id", order.id)
        span.set_attribute("payment.method", order.payment_method)
        result = payment_client.charge(order)
        span.set_attribute("payment.status", result.status)

Correlating Traces with Metrics and Logs

Traces are most powerful when correlated with the other observability pillars described in our APM overview. The trace ID serves as the correlation key:

Common Distributed Tracing Patterns

Fan-Out/Fan-In

A service calls multiple downstream services in parallel (fan-out), then aggregates their responses (fan-in). In the trace, this appears as multiple sibling child spans under a single parent. The parent span's duration includes the longest child, making it easy to identify which parallel call is the bottleneck.

Async Message Processing

For message queue-based architectures (Kafka, RabbitMQ, SQS), trace context is injected into message headers by the producer and extracted by the consumer. This creates a "link" between the producer's trace and the consumer's trace, even though the message may be processed minutes or hours later.

Retry and Circuit Breaker

When a service retries a failed request or a circuit breaker opens, each retry attempt should be a new child span with attributes indicating the retry count. This makes retry storms visible in traces and helps identify services that are causing cascading failures.

Debugging with Traces: A Workflow

  1. Symptom detection: A metric alert fires — p95 latency for /checkout exceeds 3 seconds
  2. Find exemplar traces: From the metric dashboard, click an exemplar to open a trace that exhibits the high latency
  3. Identify the bottleneck: In the waterfall view, locate the span with the longest duration — the widest bar in the chart
  4. Examine attributes: The bottleneck span's attributes reveal the cause — a slow database query, a third-party API timeout, or resource contention
  5. Correlate with logs: Search logs by the trace ID for detailed error messages, stack traces, or query plans
  6. Verify the fix: After deploying a fix, compare trace durations before and after using performance measurement tools

Key Takeaways