Kubernetes transforms the monitoring problem from tracking fixed servers to observing a dynamic system where workloads migrate across nodes, scale horizontally in seconds, and share resources through a sophisticated scheduling layer. Traditional server monitoring still matters for the underlying nodes, but container orchestration demands additional metrics, different alerting approaches, and tooling that understands the relationship between pods, deployments, services, and the cluster itself.
The Kubernetes Metrics Architecture
Kubernetes exposes metrics through a layered architecture. At the foundation, the kubelet on each node runs an embedded cAdvisor that collects resource usage data for every container. The Metrics Server aggregates these per-node metrics and serves them through the Kubernetes API, enabling features like kubectl top and the Horizontal Pod Autoscaler. For comprehensive monitoring, a full metrics pipeline — typically Prometheus with node_exporter and kube-state-metrics — collects both resource metrics and cluster state information.
Understanding which component produces which metrics prevents gaps in your monitoring coverage:
- cAdvisor (via kubelet) — container-level CPU, memory, filesystem, and network usage. These are the raw resource metrics for each running container.
- kube-state-metrics — cluster state derived from the Kubernetes API: deployment replica counts, pod phase transitions, node conditions, PersistentVolumeClaim status. These are not resource metrics but operational state.
- node_exporter — host-level metrics from the underlying node: system CPU, memory, disk I/O, network interfaces. Essential because container metrics alone cannot reveal node-level problems like kernel panics or NIC failures.
- API server metrics — request latency, etcd operation duration, admission webhook performance. These reveal control plane health.
The Metrics Server vs. Full Monitoring Pipeline
The Metrics Server provides only the current resource snapshot — no history, no alerting, no custom metrics. It serves a specific purpose: feeding real-time data to the HPA and VPA controllers and supporting kubectl top. Attempting to build a monitoring strategy solely on the Metrics Server fails because you cannot query historical data, set alert thresholds, or correlate metrics across time windows.
A production monitoring pipeline supplements the Metrics Server with Prometheus (or a compatible TSDB) that scrapes metrics at regular intervals and retains them for weeks or months. This historical data enables capacity planning, trend analysis, and anomaly detection that real-time snapshots cannot support.
Pod-Level Resource Monitoring
Containers within a pod share network and storage namespaces but maintain separate CPU and memory accounting. cAdvisor tracks each container independently, reporting metrics that include CPU usage in core-seconds, memory working set in bytes, filesystem reads and writes, and network traffic per interface.
CPU Metrics for Containers
Container CPU usage differs fundamentally from host CPU monitoring because of cgroups enforcement. When you set a CPU limit of 500m (half a core), the kernel throttles the container once it consumes 50ms of CPU time within any 100ms period. The metric container_cpu_cfs_throttled_periods_total counts how often this throttling occurs. High throttling rates indicate that the CPU limit is constraining application performance.
# Prometheus query: CPU throttling percentage per container
rate(container_cpu_cfs_throttled_periods_total{container!=""}[5m])
/
rate(container_cpu_cfs_periods_total{container!=""}[5m])
* 100
# Alert when throttling exceeds 25% of periods
# This means the container is hitting its CPU limit
# in more than one quarter of scheduling periods
CPU requests, unlike limits, do not cause throttling. Requests guarantee that the container receives at least that much CPU when contention exists but allow it to consume more when idle capacity is available. Monitor the ratio of actual usage to requests to identify containers with over-provisioned requests that waste scheduler capacity.
Memory Metrics and OOMKill Events
Memory limits in Kubernetes are hard boundaries. When a container's memory usage reaches its limit, the kernel OOM-kills the process inside the container, and Kubernetes restarts it according to the pod's restart policy. The metric kube_pod_container_status_restarts_total combined with kube_pod_container_status_last_terminated_reason="OOMKilled" reveals containers dying from memory exhaustion.
The critical memory metric is working set bytes, not RSS. Working set represents the amount of memory that cannot be reclaimed under pressure — it excludes inactive pages that the kernel would free before triggering an OOMKill. Monitor working set as a percentage of the memory limit and alert when it consistently exceeds 80%, giving the application headroom for memory spikes.
Horizontal Pod Autoscaler Monitoring
The Horizontal Pod Autoscaler (HPA) adjusts replica counts based on observed metrics, but it operates with inherent delays that monitoring must account for. The default sync period is 15 seconds, the stabilization window prevents rapid scale-down, and new pods need time to start and become ready.
HPA Metrics to Track
Monitor these HPA-specific metrics to evaluate autoscaling effectiveness:
- Current vs. desired replicas —
kube_horizontalpodautoscaler_status_current_replicascompared tokube_horizontalpodautoscaler_status_desired_replicas. A persistent gap indicates that new pods cannot schedule (insufficient node capacity) or cannot become ready (health check failures). - Current metric value vs. target — the metric driving the scaling decision compared to the configured target. If CPU target is 70% and the observed value consistently stays at 90%, the HPA is scaling too slowly or has reached its maximum replica count.
- Scaling events over time — frequent scale-up and scale-down cycles (flapping) indicate that the target metric is too close to the natural oscillation range of the workload. Widen the stabilization window or adjust the target percentage to reduce flapping.
# HPA configuration with monitoring-friendly settings
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-frontend
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web-frontend
minReplicas: 3
maxReplicas: 50
behavior:
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 10
periodSeconds: 60
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 15
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65
Cluster Health Monitoring
Cluster health extends beyond individual pod and node metrics to encompass the control plane components that coordinate the entire system.
etcd Performance
etcd is the persistence layer for all Kubernetes state. When etcd performance degrades, every cluster operation slows: pod scheduling, service updates, configmap changes. Monitor etcd write latency (etcd_disk_wal_fsync_duration_seconds) and read latency. The p99 write latency should stay below 10ms. Values above 25ms indicate storage subsystem problems that will cascade into API server timeouts and scheduling delays.
API Server Request Latency
The API server handles every interaction with the cluster. Track apiserver_request_duration_seconds bucketed by verb (GET, LIST, CREATE, DELETE) and resource type. LIST operations on large collections (pods across all namespaces) are the most common source of API server pressure, especially when monitoring tools or controllers poll aggressively.
Node Conditions
Kubernetes reports node health through conditions: Ready, MemoryPressure, DiskPressure, PIDPressure, and NetworkUnavailable. A node transitioning from Ready=True to Ready=False triggers pod eviction after the pod-eviction-timeout (default 5 minutes). Monitor node condition transitions rather than their current state — the transition signals the problem, while the current state may have already been remediated by eviction.
The kubelet also reports system logs that reveal node-level problems invisible to the Kubernetes API: kernel panics, filesystem corruption, NIC resets. Forward kubelet and kernel logs to your centralized log aggregation system alongside container logs.
Service Mesh Observability
Service meshes like Istio and Linkerd inject sidecar proxies alongside each application container, creating a uniform observability layer across all service-to-service communication. These proxies generate metrics without any instrumentation changes to application code.
Golden Signals from Sidecar Proxies
Service mesh sidecars automatically emit the four golden signals for every inter-service call:
- Latency — request duration histograms segmented by source service, destination service, and response code. This reveals which service-to-service paths contribute most to end-user latency.
- Traffic — request rate per service pair. Combined with response codes, this shows the load each service handles and whether failures correlate with traffic volume.
- Errors — the rate of responses with 5xx status codes or connection failures. Segmented by source and destination to identify which service is failing and which clients are affected.
- Saturation — the proxy's own resource usage and queue depth. When the sidecar proxy saturates, it adds latency to every request regardless of the application's health.
These proxy-generated metrics complement application-level APM data by providing consistent measurement points at service boundaries. Even services that lack internal instrumentation become observable through their mesh traffic patterns.
Monitoring Ephemeral Workloads
Jobs, CronJobs, and short-lived pods present monitoring challenges because they may complete or fail before a scrape interval captures their metrics. Use Prometheus Pushgateway for batch jobs that need to report completion metrics, or implement pod lifecycle hooks that write status to a persistent store before the pod terminates.
Track CronJob execution through kube_cronjob_status_last_schedule_time and kube_job_status_succeeded / kube_job_status_failed. A CronJob that stops scheduling — perhaps because a previous run is still active and the concurrency policy is Forbid — will not trigger a pod-level alert because no pod exists to fail. Only cluster-state metrics from kube-state-metrics capture this class of silent failure.
The interplay between container metrics, cluster state, and alerting configuration determines whether your monitoring system catches failures proactively or merely records them for postmortem analysis. Each layer provides visibility that the others cannot, and gaps at any layer create blind spots where problems incubate undetected.