Home › Observability & SRE › OpenTelemetry Implementation

OpenTelemetry: Unified Observability Framework

Observability instrumentation has historically meant choosing between incompatible vendor-specific libraries. Instrumenting an application for Datadog meant using Datadog's tracing library, which could not export to Jaeger. Switching vendors required re-instrumenting every service. OpenTelemetry eliminates this lock-in by providing a single, vendor-neutral API and SDK for generating telemetry data that can be exported to any compatible backend.

OpenTelemetry (OTel) is a CNCF project that merges the OpenTracing and OpenCensus initiatives into a unified specification. It defines APIs, SDKs, and a data collection pipeline for all three observability pillars: traces, metrics, and logs. The project produces instrumentation libraries for every major programming language and a standalone Collector component that processes and routes telemetry data.

Architecture Overview

The OpenTelemetry architecture consists of four layers, each with a distinct responsibility:

OpenTelemetry Architecture Application Layer OTel API OTel SDK Auto-Instrumentation Libraries OTLP OTel Collector Receivers Processors Exporters Jaeger / Tempo (Traces) Prometheus / Mimir (Metrics) Loki / Elasticsearch (Logs) Key Concepts API: Instrument code No-op by default, zero overhead SDK: Process telemetry Sampling, batching, export Collector: Route data Transform, filter, fan-out

The API Layer

The OTel API defines the interfaces for creating telemetry: starting spans, recording metrics, emitting log records. Crucially, the API is a no-op by default. If no SDK is configured, API calls produce no telemetry and incur near-zero overhead. This design means library authors can instrument their code with OTel API calls without forcing their users to adopt any specific observability backend.

The SDK Layer

The SDK provides the implementation behind the API interfaces. It handles span creation, metric aggregation, log processing, sampling decisions, and export to backends. Each language has its own SDK implementation that plugs into the API. The SDK is where configuration happens: which exporter to use, what sampling rate to apply, how to batch data for efficient network transfer.

The Collector

The OTel Collector is a standalone binary that receives, processes, and exports telemetry data. It decouples applications from backends. Applications send telemetry to a local Collector instance using OTLP (OpenTelemetry Protocol), and the Collector routes that data to one or more backends. This architecture provides several advantages:

Instrumentation Approaches

Auto-Instrumentation

Auto-instrumentation libraries add telemetry to common frameworks and libraries without code changes. For Java, a single agent JAR attached to the JVM instruments HTTP servers (Spring, Servlet), HTTP clients (OkHttp, Apache HttpClient), database drivers (JDBC), and message queues (Kafka, RabbitMQ) automatically.

# Java auto-instrumentation java -javaagent:opentelemetry-javaagent.jar \ -Dotel.service.name=order-service \ -Dotel.exporter.otlp.endpoint=http://collector:4317 \ -jar order-service.jar # Python auto-instrumentation opentelemetry-instrument \ --service_name order-service \ --exporter_otlp_endpoint http://collector:4317 \ python app.py # Node.js auto-instrumentation (in code) const { NodeSDK } = require('@opentelemetry/sdk-node'); const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node'); const sdk = new NodeSDK({ serviceName: 'order-service', instrumentations: [getNodeAutoInstrumentations()] }); sdk.start();

Auto-instrumentation captures the structural telemetry: HTTP request spans, database query spans, cache lookup spans. This structural data reveals where time is spent but not why it is spent there. For that, you need manual instrumentation.

Manual Instrumentation

Manual instrumentation adds business-specific telemetry that auto-instrumentation cannot infer. The OTel API provides tracer and meter instances for creating custom spans and metrics:

// Custom span for a business operation const tracer = trace.getTracer('order-service'); async function processOrder(order) { return tracer.startActiveSpan('process-order', async (span) => { span.setAttribute('order.id', order.id); span.setAttribute('order.total', order.total); span.setAttribute('order.items_count', order.items.length); try { const validated = await validateInventory(order); span.addEvent('inventory-validated', { items: validated.length }); const payment = await chargePayment(order); span.setAttribute('payment.method', payment.method); span.setAttribute('payment.processor', payment.processor); span.setStatus({ code: SpanStatusCode.OK }); return payment; } catch (error) { span.setStatus({ code: SpanStatusCode.ERROR, message: error.message }); span.recordException(error); throw error; } finally { span.end(); } }); }

The combination of auto and manual instrumentation provides both structural visibility (which services and operations are involved) and business context (what the operation means and what data it processes). Most teams start with auto-instrumentation for immediate value and add manual instrumentation incrementally for their most critical business flows.

Collector Pipeline Configuration

The Collector pipeline consists of three component types chained together:

Receivers

Receivers accept telemetry data from applications or other Collectors. The OTLP receiver handles data in OpenTelemetry's native format over gRPC or HTTP. Additional receivers support legacy formats: Jaeger, Zipkin, Prometheus, StatsD, and Fluent Forward. This allows gradual migration from existing instrumentation to OTel.

Processors

Processors transform data between receipt and export. Common processors include:

ProcessorPurposeUse Case
batchGroups telemetry into batches for efficient exportAlways use; reduces network calls
memory_limiterPrevents collector OOM by dropping data under pressureProduction safety net
attributesAdds, modifies, or removes attributesAdding environment or region labels
filterDrops telemetry matching conditionsExcluding health check spans
tail_samplingKeeps traces based on complete trace analysisAlways keep error and slow traces
transformApplies OTTL transformations to dataRedacting PII from log messages

Exporters

Exporters send processed data to backends. Each exporter targets a specific protocol or vendor: OTLP (to another Collector or OTLP-compatible backend), Prometheus (metrics), Jaeger (traces), Loki (logs), or vendor-specific exporters for commercial platforms.

A typical Collector configuration file:

receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 http: endpoint: 0.0.0.0:4318 processors: batch: timeout: 5s send_batch_size: 1024 memory_limiter: limit_mib: 512 spike_limit_mib: 128 check_interval: 5s attributes: actions: - key: environment value: production action: upsert exporters: otlp/tempo: endpoint: tempo:4317 prometheusremotewrite: endpoint: http://mimir:9009/api/v1/push loki: endpoint: http://loki:3100/loki/api/v1/push service: pipelines: traces: receivers: [otlp] processors: [memory_limiter, batch, attributes] exporters: [otlp/tempo] metrics: receivers: [otlp] processors: [memory_limiter, batch] exporters: [prometheusremotewrite] logs: receivers: [otlp] processors: [memory_limiter, batch] exporters: [loki]

Deployment Patterns

Sidecar Pattern

In Kubernetes environments, deploy a Collector as a sidecar container alongside each application pod. This provides per-pod processing and isolation. The application sends telemetry to localhost, minimizing network latency and eliminating cross-pod network failures. The sidecar pattern works well for environments with heterogeneous processing requirements where different services need different Collector configurations.

DaemonSet Pattern

Deploy a Collector as a DaemonSet with one instance per node. Applications on the same node send telemetry to the node-local Collector. This reduces resource consumption compared to sidecars (one Collector per node vs. one per pod) while maintaining data locality. The DaemonSet pattern is the most common production deployment for Kubernetes clusters.

Gateway Pattern

A centralized Collector deployment that all applications send telemetry to. The gateway handles cross-cutting processing like tail-based sampling (which requires complete traces), data enrichment from external sources, and multi-tenant routing. Gateway Collectors often sit behind a load balancer for high availability.

Production recommendation: Use a two-tier architecture. DaemonSet Collectors on each node handle initial processing, batching, and retry. A Gateway Collector handles tail-based sampling, cross-service enrichment, and export to backends. This combines the reliability of local collection with the intelligence of centralized processing.

Performance Considerations

Instrumentation overhead must be negligible to justify its presence in production. OTel is designed with performance in mind, but careless usage can introduce measurable overhead:

Migration Strategy

Migrating from existing instrumentation to OpenTelemetry does not require a big-bang switch. The Collector's multi-format receivers allow gradual migration:

  1. Deploy the Collector: Set up an OTel Collector that receives from existing formats (Jaeger, Zipkin, Prometheus) and exports to your current backends. No application changes needed.
  2. Add auto-instrumentation: Enable OTel auto-instrumentation alongside existing libraries. Both produce telemetry that the Collector merges.
  3. Migrate custom instrumentation: Replace vendor-specific manual instrumentation with OTel API calls, one service at a time.
  4. Remove legacy libraries: Once a service is fully instrumented with OTel, remove the old vendor libraries.
  5. Optimize Collector pipelines: With all telemetry flowing through OTel, optimize processing, sampling, and routing in the Collector configuration.

This incremental approach means no service experiences a disruption in observability during migration. Both old and new instrumentation coexist until the migration is complete.

Key Takeaways