Home › Observability & SRE › Metrics Cardinality Management

Metrics Cardinality Management: Controlling Observability Costs

Your Prometheus instance crashes at 3 AM. Investigation reveals that a single metric with an unbounded user_id label generated 4.2 million time series overnight. The cardinality explosion consumed all available memory, and every dashboard dependent on that instance went dark. This scenario plays out at organizations of every size. Metrics cardinality—the total number of unique time series produced by your instrumentation—is the single biggest driver of observability infrastructure costs and stability problems.

Understanding and controlling cardinality is not an optimization exercise. It is a prerequisite for running a sustainable observability practice at any scale beyond a handful of services.

What Is Cardinality?

In the context of metrics, cardinality refers to the number of unique time series generated by a single metric. A time series is defined by its metric name plus the complete set of label key-value pairs. The metric http_requests_total with labels method, status, and endpoint produces one time series for every unique combination of those label values.

With 4 HTTP methods, 10 status codes, and 50 endpoints, this metric generates 4 × 10 × 50 = 2,000 time series. That is manageable. But add a user_id label with 100,000 unique users and the number jumps to 200 million time series. That single label addition multiplied cardinality by 100,000x.

Cardinality Explosion: Label Multiplication method 4 values × status 10 values × endpoint 50 values = 2,000 time series + Add user_id label (100K users) 4 methods × 10 statuses × 50 endpoints × user_id: 100K = 200,000,000 time series (100,000x increase)

Why High Cardinality Matters

Every unique time series consumes memory in the metrics backend. Prometheus stores active time series in RAM for fast querying. Grafana Mimir, Thanos, and Cortex all maintain indexes proportional to the number of active series. The costs compound across three dimensions:

Commercial observability platforms like Datadog and New Relic price directly on custom metrics volume. An unbounded label on a high-throughput metric can generate bills in the tens of thousands of dollars per month from a single instrumentation mistake.

Common Sources of Cardinality Explosion

Unbounded Labels

Labels whose value set grows with system usage rather than system configuration are the primary source of explosions. User IDs, session tokens, request IDs, IP addresses, and full URL paths all produce one new time series per unique value. These labels should never appear on metrics.

High-Cardinality Compound Labels

Even individually bounded labels can combine to produce explosive cardinality. A metric with 6 labels, each having 20 possible values, produces 206 = 64 million potential time series. Not all combinations will appear in practice, but the theoretical maximum represents memory that the metrics backend must be prepared to handle.

Default Instrumentation Libraries

Some libraries instrument with more labels than necessary. HTTP middleware that automatically records path as a label treats every unique URL (including query parameters or path parameters like /users/12345) as a distinct value. Without path normalization, a REST API with resource IDs in paths generates unbounded cardinality.

Cardinality Reduction Strategies

Label Allowlisting

Define an explicit allowlist of label names permitted on each metric. Reject any label not on the list. This prevents accidental additions of high-cardinality labels during development. Implement the allowlist in your instrumentation SDK configuration or as a Collector processor.

Value Bucketing

Replace high-cardinality values with bounded buckets. Instead of recording exact response sizes, bucket them into ranges (0-1KB, 1-10KB, 10-100KB, 100KB+). Instead of recording full URL paths, normalize them to route patterns (/users/{id} instead of /users/12345). Histograms naturally bucket continuous values.

# Prometheus relabeling to normalize HTTP paths metric_relabel_configs: - source_labels: [path] regex: '/users/[0-9]+' target_label: path replacement: '/users/{id}' - source_labels: [path] regex: '/orders/[a-f0-9-]+' target_label: path replacement: '/orders/{id}' # Drop metrics exceeding cardinality threshold - source_labels: [__name__] regex: 'debug_.*' action: drop

Aggregation and Recording Rules

Pre-aggregate high-cardinality metrics into lower-cardinality summaries using recording rules. If you need per-endpoint latency for dashboards but not per-user latency, create a recording rule that drops the user dimension:

groups: - name: cardinality_reduction interval: 30s rules: # Aggregate per-instance metrics to per-service - record: service:http_request_duration_seconds:p99 expr: | histogram_quantile(0.99, sum by (service, method, status) ( rate(http_request_duration_seconds_bucket[5m]) ) ) # Aggregate per-endpoint to per-service for cost savings - record: service:http_requests:rate5m expr: | sum by (service, method, status) ( rate(http_requests_total[5m]) )

Recording rules trade storage for query performance and cardinality control. The pre-aggregated series exist alongside the raw series but have predictably bounded cardinality. You can set shorter retention on the high-cardinality raw data and longer retention on the aggregated summaries.

Metric Lifecycle Management

Metrics accumulate over time as teams add instrumentation but rarely remove it. Run a regular audit to identify metrics that no dashboard queries, no alert references, and no runbook mentions. These unused metrics can be dropped at the scrape level without impact. Tools like Grafana Mimirtool can analyze ruler and dashboard configurations to identify truly unused metrics.

Cardinality Budgets

Assign a cardinality budget to each team or service. A service producing 50,000 active time series is healthy. A service producing 5 million active series needs investigation. Set budgets based on the metric backend's capacity divided among teams, with headroom for growth.

Service TierSeries BudgetAlert ThresholdAction
Small (1-5 pods)10,0008,000Review labels
Medium (5-20 pods)50,00040,000Optimize recording rules
Large (20-100 pods)200,000160,000Mandatory aggregation
Critical (100+ pods)500,000400,000Dedicated review + federation

Monitor actual cardinality against budgets using Prometheus's own metrics. The prometheus_tsdb_head_series metric shows total active series. For per-metric breakdowns, scrape_series_added reveals which targets contribute the most series.

Warning: The count({__name__=~".+"}) query itself is expensive on high-cardinality instances. Use TSDB status API endpoints (/api/v1/status/tsdb) for cardinality analysis instead of brute-force queries.

Implementation Patterns

Gate at Ingestion

The most effective control point is at data ingestion, before series reach the backend. Configure the OTel Collector or Prometheus scrape configuration to drop or relabel problematic metrics before they consume resources.

Alert on Cardinality Growth

Set alerts for sudden increases in active series counts. A 20% increase in active series within an hour usually indicates a deployment introduced a new high-cardinality label or a new unbounded metric.

# Alert on sudden cardinality increase - alert: HighCardinalityGrowth expr: | (prometheus_tsdb_head_series - prometheus_tsdb_head_series offset 1h) / prometheus_tsdb_head_series offset 1h > 0.2 for: 15m labels: severity: warning annotations: summary: "Active series count grew >20% in the last hour"

CI/CD Integration

Shift cardinality checks left into the development pipeline. Lint instrumentation code to detect unbounded labels before they reach production. Tools like promtool check metrics can validate metric naming and label conventions. Code review checklists should include cardinality impact assessment for any new metric or label.

Measuring Success

Track these SLI-style metrics for your cardinality management program:

Key Takeaways