Home›Server & Infrastructure Monitoring›Monitoring IaC Deployments
Server & Infrastructure Monitoring

Monitoring Infrastructure as Code Deployments

Infrastructure as Code transforms infrastructure management from manual console clicks into version-controlled, reviewable, repeatable processes. But this shift creates a new monitoring challenge: tracking whether the declared state in your Terraform files, Pulumi programs, or CloudFormation templates actually matches what is running in production. Drift detection, deployment health verification, and canary analysis become essential observability layers that protect the promise IaC makes — that the code is the truth.

The Deployment Monitoring Gap

Most infrastructure monitoring focuses on steady-state operations: CPU, memory, and network metrics collected continuously from running systems. The gap appears during state transitions — when infrastructure is being provisioned, updated, or decommissioned. A Terraform apply that creates 47 resources can fail after resource 23, leaving infrastructure in a partially provisioned state that neither the old configuration nor the new configuration accurately describes.

Monitoring this transition period requires visibility into:

  • Plan accuracy — did the plan predict the right number of creates, updates, and destroys?
  • Apply progress — how many resources have been provisioned, and which ones failed?
  • Post-apply verification — are the provisioned resources actually functional, not just present?
  • Drift over time — has the running infrastructure diverged from the declared state since the last apply?

Terraform State Monitoring

Terraform state is the record of what Terraform believes exists in the target environment. State corruption, staleness, or lock contention cause deployment failures that manifest as mysterious errors during plan or apply operations.

State File Health Checks

Monitor the Terraform state backend for availability and integrity. For S3-backed state:

# State file health monitoring
# Check state file exists and is recent
aws s3api head-object \
  --bucket terraform-state-prod \
  --key infrastructure/terraform.tfstate \
  --query 'LastModified'

# Check state lock table is accessible
aws dynamodb describe-table \
  --table-name terraform-locks \
  --query 'Table.TableStatus'

# Verify state file is valid JSON
terraform show -json | jq '.format_version' 2>/dev/null
# Returns version string if valid, error if corrupt

State lock monitoring prevents the scenario where a CI/CD pipeline holds a lock indefinitely — perhaps due to a crashed runner or network partition — blocking all subsequent deployments. Monitor the DynamoDB lock table for entries older than your maximum expected apply duration (typically 30-60 minutes) and alert when stale locks are detected.

State Drift Detection

Drift occurs when the real infrastructure diverges from the Terraform state. Manual changes through the console, automated scripts that modify resources directly, or external systems that update configurations all create drift. Running terraform plan periodically reveals drift as unexpected diff output.

# Automated drift detection pipeline
# Run daily via CI/CD schedule

# 1. Generate plan in machine-readable format
terraform plan -detailed-exitcode -out=drift.plan 2>&1
EXIT_CODE=$?

# Exit codes:
# 0 = no changes (no drift)
# 1 = error
# 2 = changes detected (drift found)

if [ $EXIT_CODE -eq 2 ]; then
  # Extract drift summary
  terraform show -json drift.plan | jq '{
    drift_resources: [
      .resource_changes[]
      | select(.change.actions != ["no-op"])
      | {address, actions: .change.actions}
    ],
    total_drift: [
      .resource_changes[]
      | select(.change.actions != ["no-op"])
    ] | length
  }'

  # Send drift alert
  # Include resource addresses and change types
fi
IaC Deployment Monitoring Pipeline Plan Diff detection Cost estimation Apply Resource creation Progress tracking Verify Health checks Smoke tests Monitor Drift detection Compliance audit Rollback on failure Plan duration / changes Apply success rate Post-deploy health Drift count / severity

Canary Deployment Analysis

Canary deployments route a small percentage of traffic to newly provisioned infrastructure while maintaining the existing version for the majority of users. The monitoring challenge is determining whether the canary performs acceptably before promoting it to handle all traffic.

Canary Metrics Comparison

Compare the canary against the baseline across four dimensions:

  • Error rate — the canary's 5xx rate compared to the baseline's 5xx rate. A statistically significant increase indicates the new infrastructure introduces errors.
  • Latency distribution — compare p50, p95, and p99 latency between canary and baseline. The canary should match or improve on baseline latency at all percentiles.
  • Resource consumption — CPU and memory usage on the canary relative to baseline. Higher resource consumption may indicate inefficiency in the new configuration.
  • Business metrics — conversion rate, throughput, or other application-specific KPIs. Infrastructure changes that degrade business outcomes should fail the canary even if technical metrics appear healthy.

Automated Rollback Triggers

Define quantitative criteria that automatically roll back a canary deployment:

# Canary analysis configuration
canary:
  analysis:
    duration: 30m
    interval: 1m
    thresholds:
      # Roll back if canary error rate exceeds baseline by 0.5%
      error_rate_increase: 0.005
      # Roll back if canary p99 latency exceeds baseline by 20%
      p99_latency_increase_pct: 20
      # Roll back if canary p50 latency exceeds baseline by 10%
      p50_latency_increase_pct: 10
      # Minimum traffic for statistical significance
      min_requests: 1000
  rollback:
    automatic: true
    # Require 3 consecutive failing checks before rollback
    failure_threshold: 3

Deployment Frequency and Change Failure Rate

Two of the four DORA metrics — deployment frequency and change failure rate — directly relate to IaC monitoring. Tracking these metrics over time reveals whether your infrastructure delivery practice is improving or regressing.

Deployment frequency measures how often you successfully deploy infrastructure changes to production. Higher frequency indicates smaller, less risky changes and greater confidence in the deployment pipeline. Track this by counting successful terraform apply runs in production per week.

Change failure rate measures the percentage of deployments that result in degraded service, require rollback, or need hotfix. Calculate this from your deployment records and alert correlation: of the 47 deployments this month, how many triggered a SEV 1 or SEV 2 alert within 30 minutes of completion?

Compliance and Policy Monitoring

IaC enables policy-as-code enforcement through tools like Open Policy Agent (OPA), Sentinel, and Conftest. These tools evaluate Terraform plans against organizational policies before apply, preventing non-compliant infrastructure from being provisioned.

Policy Violations as Metrics

Track policy violations as monitoring metrics rather than just CI/CD gate failures. The trend of policy violations over time reveals whether teams are learning from enforcement or repeatedly attempting non-compliant configurations. A rising violation count in a specific policy category might indicate that the policy is misaligned with a legitimate engineering need, warranting policy revision rather than continued enforcement.

# OPA policy: require encryption on all S3 buckets
package terraform.aws.s3

deny[msg] {
  resource := input.resource_changes[_]
  resource.type == "aws_s3_bucket"
  resource.change.after.server_side_encryption_configuration == null
  msg := sprintf(
    "S3 bucket '%s' must have server-side encryption enabled",
    [resource.address]
  )
}

# Track violations as custom metrics
# violation_category: encryption
# violation_resource: aws_s3_bucket
# violation_team: derived from workspace/path

Cost Monitoring for Infrastructure Changes

Infrastructure cost estimation before deployment prevents budget surprises. Tools like Infracost analyze Terraform plans and estimate the monthly cost impact of proposed changes. Integrating cost estimation into the deployment pipeline makes cost a first-class deployment metric alongside performance and reliability.

Monitor cost drift alongside infrastructure drift. A resource that was provisioned as a t3.medium instance but manually upgraded to a c5.2xlarge through the console represents both configuration drift and cost drift. The log trail of console actions, combined with periodic plan comparisons, reveals these unauthorized cost increases.

Set cost change thresholds that require additional approval. Any deployment that increases monthly infrastructure cost by more than $500 or 10% triggers a review by the engineering manager. This threshold should be codified in the CI/CD pipeline, not left to manual review of plan output.

Post-Deployment Health Verification

A successful terraform apply means resources exist in the target environment, not that they function correctly. Post-deployment health checks validate that provisioned infrastructure actually serves its purpose.

Implement health verification as a pipeline stage that runs after apply completes:

  1. Connectivity checks — can dependent services reach the new infrastructure? Test network paths, DNS resolution, and TLS handshakes.
  2. Functional checks — does the infrastructure perform its intended function? For a database, verify that connections succeed and queries return results. For a load balancer, verify that health checks pass for backend targets.
  3. Performance checks — does the infrastructure meet performance expectations? Run a brief synthetic monitoring check against the new resources to verify latency and throughput.

Failure at any verification stage should trigger an automatic rollback to the previous infrastructure version, returning the environment to a known-good state while the team investigates the failure cause. The verification results become part of the deployment record, providing evidence that each infrastructure change was validated before being declared complete.