Home›APM›Error Tracking

Error Tracking and Monitoring: From Detection to Resolution

Carlos ReyesSeptember 18, 202613 min read

Every production system generates errors. The difference between a well-operated service and a fragile one is not the absence of errors — it is the speed and precision with which errors are detected, categorized, triaged, and resolved. Error tracking transforms the raw stream of exceptions, failed requests, and unexpected behaviors into actionable intelligence that teams can act on before users complain.

This article covers the complete error tracking lifecycle: from capturing and categorizing errors, through deduplication and grouping, to error budgets, alerting thresholds, and incident response integration. For context on how error tracking fits within the broader observability stack, see our APM fundamentals guide.

Error Taxonomy

Not all errors are equal. A robust error tracking system categorizes errors by their source, severity, and impact to enable effective prioritization.

Production Errors Application Errors Unhandled exceptions Logic errors Validation failures Timeout errors Infrastructure Errors OOM kills Disk full Network partitions DNS failures Dependency Errors API failures (3rd party) DB connection lost Queue backpressure Cache misses (bulk) Error Classification Hierarchy

Error Severity Levels

SeverityDefinitionResponseExample
Critical (P0)Complete service outage or data lossImmediate page, all-hands responseDatabase corruption, auth service down
High (P1)Major feature broken, significant user impactPage on-call, fix within hoursPayment processing failures, search returning empty
Medium (P2)Feature degraded, workaround existsQueue for next business daySlow image loading, non-critical API timeouts
Low (P3)Minor issue, cosmetic, edge caseBacklog for sprint planningTooltip rendering glitch, rare input validation miss

Error Capture and Enrichment

Raw errors are often insufficient for debugging. Effective error tracking captures not just the exception type and message, but also the surrounding context that makes reproduction possible.

Essential Error Context

# Python: Structured error capture with context
import logging
import traceback
from contextvars import ContextVar

request_id_var = ContextVar('request_id', default='unknown')

class StructuredErrorHandler:
    def capture(self, error, context=None):
        error_event = {
            "type": type(error).__name__,
            "message": str(error),
            "stack_trace": traceback.format_exception(error),
            "fingerprint": self._compute_fingerprint(error),
            "request_id": request_id_var.get(),
            "timestamp": datetime.utcnow().isoformat(),
            "severity": self._classify_severity(error),
            "tags": {
                "service": "order-service",
                "version": os.getenv("APP_VERSION", "unknown"),
                "region": os.getenv("AWS_REGION", "unknown"),
            },
            "context": context or {},
        }
        self._send_to_backend(error_event)

    def _compute_fingerprint(self, error):
        """Group identical errors by type + location"""
        frames = traceback.extract_tb(error.__traceback__)
        if frames:
            last_frame = frames[-1]
            return f"{type(error).__name__}:{last_frame.filename}:{last_frame.lineno}"
        return f"{type(error).__name__}:unknown"

    def _classify_severity(self, error):
        if isinstance(error, (ConnectionError, TimeoutError)):
            return "high"
        if isinstance(error, (ValueError, KeyError)):
            return "medium"
        return "low"

Error Grouping and Deduplication

A single bug can generate thousands of identical error events per minute. Without grouping, the error tracking system becomes a firehose of noise. Effective grouping collapses identical or similar errors into "issues" — each issue represents one underlying bug, regardless of how many times it fires.

Fingerprinting Strategies

Grouping accuracy is the most important feature of an error tracking tool. Over-grouping (merging distinct bugs into one issue) hides problems. Under-grouping (splitting one bug into many issues) creates noise. Review your top-volume error groups monthly — if a single issue contains errors from unrelated code paths, refine the fingerprinting.

Error Budgets and SLOs

An error budget defines how many errors are acceptable before action is required. It operationalizes the Service Level Objective (SLO) — the reliability target — by converting it into a concrete error allowance.

Calculating Error Budgets

If your SLO states that 99.9% of requests must succeed, then your error budget is 0.1% of total requests. For a service handling 1 million requests per day, the error budget is 1,000 errors per day.

# Error budget calculation
slo_target = 0.999          # 99.9% success rate
daily_requests = 1_000_000

error_budget = daily_requests * (1 - slo_target)
# error_budget = 1000 errors/day

# Current burn rate
current_errors_today = 450
budget_consumed = current_errors_today / error_budget
# budget_consumed = 0.45 (45% of daily budget used)

# Projected budget consumption
hours_elapsed = 8
projected_daily = (current_errors_today / hours_elapsed) * 24
# projected_daily = 1350 errors → exceeds budget

# Alert thresholds (burn rate multipliers)
# 1-hour burn rate: if last hour consumed >2% of daily budget → warning
# 6-hour burn rate: if last 6 hours consumed >5% of daily budget → alert
# Alert when projected to exhaust budget before end of window

Burn Rate Alerts

Rather than alerting on individual error counts, burn rate alerts measure how quickly the error budget is being consumed. A burn rate of 1x means the budget will be exactly exhausted by the end of the window. A burn rate of 10x means the budget will be consumed in one-tenth of the window — a clear emergency.

Burn RateBudget ExhaustionAlert SeverityResponse
14.4x~1 hourCritical (page)Immediate incident response
6x~4 hoursHigh (page)Triage within 30 minutes
3x~8 hoursWarningInvestigate during business hours
1x~24 hoursInfoReview during sprint planning

Error Monitoring in the Frontend

Frontend errors present unique challenges compared to backend errors. JavaScript exceptions occur in diverse browser environments, network conditions, and device capabilities. An error that appears on one browser version may not exist on another.

Key Frontend Error Types

Source Map Integration

Production JavaScript is minified and bundled, making raw stack traces unreadable. Source maps translate minified positions back to original file names and line numbers. Upload source maps to your error tracking service as part of the deployment pipeline — never serve source maps to end users.

Incident Response Integration

Error tracking feeds directly into incident response. When an error group crosses its severity threshold, the system should automatically:

  1. Create an incident: Open a ticket in the incident management system (PagerDuty, OpsGenie, or equivalent) with the error group details, affected users count, and relevant dashboards
  2. Page the on-call: Route the notification to the team that owns the affected service. Ownership metadata on the error group determines routing.
  3. Assemble context: Pull together the error stack trace, recent deployments, related trace samples, and infrastructure metrics into a single incident view
  4. Track resolution: Link the incident to the pull request that fixes it, and verify error rates return to baseline after deployment

Error Tracking Anti-Patterns

Building an Error Tracking Pipeline

# Error tracking pipeline architecture (conceptual)

# 1. Capture: SDK in each service
capture_layer:
  - language_sdks: [python, node, java, go]
  - auto_instrumentation: true
  - breadcrumbs: true
  - context_enrichment:
      - request_id
      - trace_id
      - user_id (hashed)
      - deployment_version

# 2. Transport: Batched, async, with backpressure
transport_layer:
  - protocol: HTTPS POST
  - batching: 100 events or 5 seconds
  - retry: exponential backoff, max 3 attempts
  - sampling: keep all errors, sample info-level events

# 3. Processing: Fingerprint, deduplicate, enrich
processing_layer:
  - fingerprinting: stack_trace + exception_type
  - deduplication: merge events with same fingerprint
  - enrichment:
      - source_map_lookup
      - git_blame (map stack frames to authors)
      - release_correlation

# 4. Storage and alerting
storage_layer:
  - event_store: 90 days retention
  - issue_store: indefinite (aggregated)
  - alerting:
      - burn_rate_thresholds: [14.4x, 6x, 3x]
      - new_error_type: immediate notification
      - regression: alert when resolved issue reappears

Key Takeaways