OpenTelemetry: Unified Observability Framework
Observability instrumentation has historically meant choosing between incompatible vendor-specific libraries. Instrumenting an application for Datadog meant using Datadog's tracing library, which could not export to Jaeger. Switching vendors required re-instrumenting every service. OpenTelemetry eliminates this lock-in by providing a single, vendor-neutral API and SDK for generating telemetry data that can be exported to any compatible backend.
OpenTelemetry (OTel) is a CNCF project that merges the OpenTracing and OpenCensus initiatives into a unified specification. It defines APIs, SDKs, and a data collection pipeline for all three observability pillars: traces, metrics, and logs. The project produces instrumentation libraries for every major programming language and a standalone Collector component that processes and routes telemetry data.
Architecture Overview
The OpenTelemetry architecture consists of four layers, each with a distinct responsibility:
The API Layer
The OTel API defines the interfaces for creating telemetry: starting spans, recording metrics, emitting log records. Crucially, the API is a no-op by default. If no SDK is configured, API calls produce no telemetry and incur near-zero overhead. This design means library authors can instrument their code with OTel API calls without forcing their users to adopt any specific observability backend.
The SDK Layer
The SDK provides the implementation behind the API interfaces. It handles span creation, metric aggregation, log processing, sampling decisions, and export to backends. Each language has its own SDK implementation that plugs into the API. The SDK is where configuration happens: which exporter to use, what sampling rate to apply, how to batch data for efficient network transfer.
The Collector
The OTel Collector is a standalone binary that receives, processes, and exports telemetry data. It decouples applications from backends. Applications send telemetry to a local Collector instance using OTLP (OpenTelemetry Protocol), and the Collector routes that data to one or more backends. This architecture provides several advantages:
- Backend portability: Switching from Jaeger to Tempo requires changing the Collector configuration, not the application code.
- Data processing: The Collector can filter sensitive data, add metadata, transform formats, and sample before export.
- Fan-out: The same telemetry data can be sent to multiple backends simultaneously for different purposes.
- Reliability: The Collector buffers data and retries failed exports, protecting against transient backend outages.
Instrumentation Approaches
Auto-Instrumentation
Auto-instrumentation libraries add telemetry to common frameworks and libraries without code changes. For Java, a single agent JAR attached to the JVM instruments HTTP servers (Spring, Servlet), HTTP clients (OkHttp, Apache HttpClient), database drivers (JDBC), and message queues (Kafka, RabbitMQ) automatically.
Auto-instrumentation captures the structural telemetry: HTTP request spans, database query spans, cache lookup spans. This structural data reveals where time is spent but not why it is spent there. For that, you need manual instrumentation.
Manual Instrumentation
Manual instrumentation adds business-specific telemetry that auto-instrumentation cannot infer. The OTel API provides tracer and meter instances for creating custom spans and metrics:
The combination of auto and manual instrumentation provides both structural visibility (which services and operations are involved) and business context (what the operation means and what data it processes). Most teams start with auto-instrumentation for immediate value and add manual instrumentation incrementally for their most critical business flows.
Collector Pipeline Configuration
The Collector pipeline consists of three component types chained together:
Receivers
Receivers accept telemetry data from applications or other Collectors. The OTLP receiver handles data in OpenTelemetry's native format over gRPC or HTTP. Additional receivers support legacy formats: Jaeger, Zipkin, Prometheus, StatsD, and Fluent Forward. This allows gradual migration from existing instrumentation to OTel.
Processors
Processors transform data between receipt and export. Common processors include:
| Processor | Purpose | Use Case |
|---|---|---|
| batch | Groups telemetry into batches for efficient export | Always use; reduces network calls |
| memory_limiter | Prevents collector OOM by dropping data under pressure | Production safety net |
| attributes | Adds, modifies, or removes attributes | Adding environment or region labels |
| filter | Drops telemetry matching conditions | Excluding health check spans |
| tail_sampling | Keeps traces based on complete trace analysis | Always keep error and slow traces |
| transform | Applies OTTL transformations to data | Redacting PII from log messages |
Exporters
Exporters send processed data to backends. Each exporter targets a specific protocol or vendor: OTLP (to another Collector or OTLP-compatible backend), Prometheus (metrics), Jaeger (traces), Loki (logs), or vendor-specific exporters for commercial platforms.
A typical Collector configuration file:
Deployment Patterns
Sidecar Pattern
In Kubernetes environments, deploy a Collector as a sidecar container alongside each application pod. This provides per-pod processing and isolation. The application sends telemetry to localhost, minimizing network latency and eliminating cross-pod network failures. The sidecar pattern works well for environments with heterogeneous processing requirements where different services need different Collector configurations.
DaemonSet Pattern
Deploy a Collector as a DaemonSet with one instance per node. Applications on the same node send telemetry to the node-local Collector. This reduces resource consumption compared to sidecars (one Collector per node vs. one per pod) while maintaining data locality. The DaemonSet pattern is the most common production deployment for Kubernetes clusters.
Gateway Pattern
A centralized Collector deployment that all applications send telemetry to. The gateway handles cross-cutting processing like tail-based sampling (which requires complete traces), data enrichment from external sources, and multi-tenant routing. Gateway Collectors often sit behind a load balancer for high availability.
Performance Considerations
Instrumentation overhead must be negligible to justify its presence in production. OTel is designed with performance in mind, but careless usage can introduce measurable overhead:
- Span creation cost: Creating a span allocates memory for attributes, events, and timing data. In hot loops processing millions of iterations per second, creating a span per iteration adds significant overhead. Instrument at the operation level (HTTP request, database query), not at the function level.
- Attribute cardinality: High-cardinality attributes on spans and metrics increase memory usage in the SDK's internal buffers. Apply the same cardinality management principles to span attributes as to metric labels.
- Export batching: The batch processor is essential. Without batching, every span triggers a network call to the Collector. With batching, hundreds of spans ship in a single request. Configure batch size and timeout to balance latency and efficiency.
- Sampling: Head-based sampling at rates between 1% and 10% is appropriate for high-throughput services. Combined with priority sampling for errors and slow requests, this captures the diagnostically valuable traces while keeping overhead minimal.
Migration Strategy
Migrating from existing instrumentation to OpenTelemetry does not require a big-bang switch. The Collector's multi-format receivers allow gradual migration:
- Deploy the Collector: Set up an OTel Collector that receives from existing formats (Jaeger, Zipkin, Prometheus) and exports to your current backends. No application changes needed.
- Add auto-instrumentation: Enable OTel auto-instrumentation alongside existing libraries. Both produce telemetry that the Collector merges.
- Migrate custom instrumentation: Replace vendor-specific manual instrumentation with OTel API calls, one service at a time.
- Remove legacy libraries: Once a service is fully instrumented with OTel, remove the old vendor libraries.
- Optimize Collector pipelines: With all telemetry flowing through OTel, optimize processing, sampling, and routing in the Collector configuration.
This incremental approach means no service experiences a disruption in observability during migration. Both old and new instrumentation coexist until the migration is complete.
Key Takeaways
- OpenTelemetry provides a vendor-neutral API, SDK, and Collector for producing and routing all three telemetry types: traces, metrics, and logs.
- The API is no-op by default, enabling library authors to instrument code without imposing observability dependencies on users.
- Auto-instrumentation covers structural telemetry with zero code changes. Add manual instrumentation for business-specific context.
- The Collector decouples applications from backends. Changing your observability vendor requires only a Collector configuration change.
- Deploy Collectors in a two-tier architecture: DaemonSet for local collection and reliability, Gateway for intelligent processing like tail-based sampling.