Error Tracking and Monitoring: From Detection to Resolution
Every production system generates errors. The difference between a well-operated service and a fragile one is not the absence of errors — it is the speed and precision with which errors are detected, categorized, triaged, and resolved. Error tracking transforms the raw stream of exceptions, failed requests, and unexpected behaviors into actionable intelligence that teams can act on before users complain.
This article covers the complete error tracking lifecycle: from capturing and categorizing errors, through deduplication and grouping, to error budgets, alerting thresholds, and incident response integration. For context on how error tracking fits within the broader observability stack, see our APM fundamentals guide.
Error Taxonomy
Not all errors are equal. A robust error tracking system categorizes errors by their source, severity, and impact to enable effective prioritization.
Error Severity Levels
| Severity | Definition | Response | Example |
|---|---|---|---|
| Critical (P0) | Complete service outage or data loss | Immediate page, all-hands response | Database corruption, auth service down |
| High (P1) | Major feature broken, significant user impact | Page on-call, fix within hours | Payment processing failures, search returning empty |
| Medium (P2) | Feature degraded, workaround exists | Queue for next business day | Slow image loading, non-critical API timeouts |
| Low (P3) | Minor issue, cosmetic, edge case | Backlog for sprint planning | Tooltip rendering glitch, rare input validation miss |
Error Capture and Enrichment
Raw errors are often insufficient for debugging. Effective error tracking captures not just the exception type and message, but also the surrounding context that makes reproduction possible.
Essential Error Context
- Stack trace: The call chain leading to the error, with source maps applied for minified code
- Request context: HTTP method, URL, headers (sanitized), request body (redacted)
- User context: User ID, session ID, account tier (never PII like email or name in error payloads)
- Environment: Service version, deployment ID, region, container/pod ID
- Breadcrumbs: A trail of recent events (UI clicks, API calls, state changes) leading up to the error
- Trace ID: Linking the error to its distributed trace for cross-service debugging
# Python: Structured error capture with context
import logging
import traceback
from contextvars import ContextVar
request_id_var = ContextVar('request_id', default='unknown')
class StructuredErrorHandler:
def capture(self, error, context=None):
error_event = {
"type": type(error).__name__,
"message": str(error),
"stack_trace": traceback.format_exception(error),
"fingerprint": self._compute_fingerprint(error),
"request_id": request_id_var.get(),
"timestamp": datetime.utcnow().isoformat(),
"severity": self._classify_severity(error),
"tags": {
"service": "order-service",
"version": os.getenv("APP_VERSION", "unknown"),
"region": os.getenv("AWS_REGION", "unknown"),
},
"context": context or {},
}
self._send_to_backend(error_event)
def _compute_fingerprint(self, error):
"""Group identical errors by type + location"""
frames = traceback.extract_tb(error.__traceback__)
if frames:
last_frame = frames[-1]
return f"{type(error).__name__}:{last_frame.filename}:{last_frame.lineno}"
return f"{type(error).__name__}:unknown"
def _classify_severity(self, error):
if isinstance(error, (ConnectionError, TimeoutError)):
return "high"
if isinstance(error, (ValueError, KeyError)):
return "medium"
return "low"
Error Grouping and Deduplication
A single bug can generate thousands of identical error events per minute. Without grouping, the error tracking system becomes a firehose of noise. Effective grouping collapses identical or similar errors into "issues" — each issue represents one underlying bug, regardless of how many times it fires.
Fingerprinting Strategies
- Stack trace grouping: The most common approach. Errors with the same exception type and the same top N frames in the stack trace are grouped together. Works well for deterministic errors.
- Message-based grouping: For errors with dynamic messages (e.g., "User 12345 not found"), strip variable parts and group by the template. Regular expressions or tokenization extract the static pattern.
- Custom fingerprints: Application code can explicitly set a grouping key when the default algorithm would incorrectly split or merge errors. This is essential for framework-generated errors that share stack traces but represent different bugs.
Grouping accuracy is the most important feature of an error tracking tool. Over-grouping (merging distinct bugs into one issue) hides problems. Under-grouping (splitting one bug into many issues) creates noise. Review your top-volume error groups monthly — if a single issue contains errors from unrelated code paths, refine the fingerprinting.
Error Budgets and SLOs
An error budget defines how many errors are acceptable before action is required. It operationalizes the Service Level Objective (SLO) — the reliability target — by converting it into a concrete error allowance.
Calculating Error Budgets
If your SLO states that 99.9% of requests must succeed, then your error budget is 0.1% of total requests. For a service handling 1 million requests per day, the error budget is 1,000 errors per day.
# Error budget calculation
slo_target = 0.999 # 99.9% success rate
daily_requests = 1_000_000
error_budget = daily_requests * (1 - slo_target)
# error_budget = 1000 errors/day
# Current burn rate
current_errors_today = 450
budget_consumed = current_errors_today / error_budget
# budget_consumed = 0.45 (45% of daily budget used)
# Projected budget consumption
hours_elapsed = 8
projected_daily = (current_errors_today / hours_elapsed) * 24
# projected_daily = 1350 errors → exceeds budget
# Alert thresholds (burn rate multipliers)
# 1-hour burn rate: if last hour consumed >2% of daily budget → warning
# 6-hour burn rate: if last 6 hours consumed >5% of daily budget → alert
# Alert when projected to exhaust budget before end of window
Burn Rate Alerts
Rather than alerting on individual error counts, burn rate alerts measure how quickly the error budget is being consumed. A burn rate of 1x means the budget will be exactly exhausted by the end of the window. A burn rate of 10x means the budget will be consumed in one-tenth of the window — a clear emergency.
| Burn Rate | Budget Exhaustion | Alert Severity | Response |
|---|---|---|---|
| 14.4x | ~1 hour | Critical (page) | Immediate incident response |
| 6x | ~4 hours | High (page) | Triage within 30 minutes |
| 3x | ~8 hours | Warning | Investigate during business hours |
| 1x | ~24 hours | Info | Review during sprint planning |
Error Monitoring in the Frontend
Frontend errors present unique challenges compared to backend errors. JavaScript exceptions occur in diverse browser environments, network conditions, and device capabilities. An error that appears on one browser version may not exist on another.
Key Frontend Error Types
- JavaScript exceptions: Runtime errors caught by
window.onerrororwindow.addEventListener('unhandledrejection') - Network errors: Failed API calls, CORS violations, resource loading failures
- Performance errors: Operations that exceed acceptable thresholds — long input delays, layout shifts, or unresponsive UI
- Console errors: Framework warnings (React, Vue) that indicate potential issues before they become visible bugs
Source Map Integration
Production JavaScript is minified and bundled, making raw stack traces unreadable. Source maps translate minified positions back to original file names and line numbers. Upload source maps to your error tracking service as part of the deployment pipeline — never serve source maps to end users.
Incident Response Integration
Error tracking feeds directly into incident response. When an error group crosses its severity threshold, the system should automatically:
- Create an incident: Open a ticket in the incident management system (PagerDuty, OpsGenie, or equivalent) with the error group details, affected users count, and relevant dashboards
- Page the on-call: Route the notification to the team that owns the affected service. Ownership metadata on the error group determines routing.
- Assemble context: Pull together the error stack trace, recent deployments, related trace samples, and infrastructure metrics into a single incident view
- Track resolution: Link the incident to the pull request that fixes it, and verify error rates return to baseline after deployment
Error Tracking Anti-Patterns
- Catching and swallowing exceptions:
try/catchblocks that log a generic message and continue hide real problems. Only catch exceptions you can meaningfully handle; let the rest propagate to the error tracker. - Alerting on every error: Not every error is actionable. Background job retries, expected validation failures, and bot-generated 404s are noise, not signal. Use allowlists or severity filters to suppress known non-actionable errors.
- Ignoring error volume trends: A steady baseline of 50 errors/hour may be acceptable. A jump from 50 to 200 — even if no individual error looks critical — signals a systemic issue.
- No ownership on error groups: Unowned errors stay unresolved. Every error group should have a team owner, and unassigned errors should be reviewed in a weekly triage rotation.
Building an Error Tracking Pipeline
# Error tracking pipeline architecture (conceptual)
# 1. Capture: SDK in each service
capture_layer:
- language_sdks: [python, node, java, go]
- auto_instrumentation: true
- breadcrumbs: true
- context_enrichment:
- request_id
- trace_id
- user_id (hashed)
- deployment_version
# 2. Transport: Batched, async, with backpressure
transport_layer:
- protocol: HTTPS POST
- batching: 100 events or 5 seconds
- retry: exponential backoff, max 3 attempts
- sampling: keep all errors, sample info-level events
# 3. Processing: Fingerprint, deduplicate, enrich
processing_layer:
- fingerprinting: stack_trace + exception_type
- deduplication: merge events with same fingerprint
- enrichment:
- source_map_lookup
- git_blame (map stack frames to authors)
- release_correlation
# 4. Storage and alerting
storage_layer:
- event_store: 90 days retention
- issue_store: indefinite (aggregated)
- alerting:
- burn_rate_thresholds: [14.4x, 6x, 3x]
- new_error_type: immediate notification
- regression: alert when resolved issue reappears
Key Takeaways
- Categorize errors by source (application, infrastructure, dependency) and severity (P0-P3)
- Capture rich context with every error: stack trace, request data, breadcrumbs, and trace ID
- Fingerprinting accuracy determines whether error tracking is useful or noisy
- Use error budgets and burn rate alerts instead of alerting on individual error counts
- Frontend errors need source map integration and browser-aware deduplication
- Integrate error tracking with incident response: auto-create incidents, page on-call, track resolution
- Every error group needs a team owner — unowned errors are unresolved errors