AWS CloudWatch Alarms, Logs and Dashboards with Terraform

· 1 min read · Terraform & OpenTofu Tutorials

Observability as code #

Amazon CloudWatch collects metrics, logs and events from AWS. Defining alarms and dashboards in Terraform means that every new resource ships with its monitoring and that it is reviewed like any other change.

An SNS topic to notify #

Alarms need somewhere to send the notification. Create a topic and subscribe the operations team (see SQS and SNS):

alerts.tf
resource "aws_sns_topic" "alerts" {
  name = "ditwl-pro-alerts"
}

resource "aws_sns_topic_subscription" "ops" {
  topic_arn = aws_sns_topic.alerts.arn
  protocol  = "email"
  endpoint  = "ops@example.com"
}

Confirm the subscription with the link that arrives by email.

Log groups with retention #

By default, log groups keep logs forever and cost grows silently. Always set a retention:

logs.tf
resource "aws_cloudwatch_log_group" "app" {
  name              = "/ditwl/pro/app"
  retention_in_days = 30
  kms_key_id        = aws_kms_key.logs.arn   # optional
}

An alarm for high CPU on EC2 #

alarms.tf
resource "aws_cloudwatch_metric_alarm" "ec2_cpu" {
  alarm_name          = "ditwl-pro-web-cpu-high"
  alarm_description   = "CPU above 80% for 10 minutes"
  namespace           = "AWS/EC2"
  metric_name         = "CPUUtilization"
  statistic           = "Average"
  period              = 300
  evaluation_periods  = 2
  threshold           = 80
  comparison_operator = "GreaterThanThreshold"
  treat_missing_data  = "notBreaching"

  dimensions = {
    InstanceId = aws_instance.web.id
  }

  alarm_actions = [aws_sns_topic.alerts.arn]
  ok_actions    = [aws_sns_topic.alerts.arn]
}

period is the length of each data point in seconds, and the alarm goes to ALARM when evaluation_periods consecutive points are above the threshold. treat_missing_data decides what happens when there is no data, which is important for stopped instances.

Alarms for several resources with for_each #

alarms.tf
locals {
  rds_alarms = {
    cpu = {
      metric    = "CPUUtilization"
      threshold = 80
      operator  = "GreaterThanThreshold"
    }
    free_storage = {
      metric    = "FreeStorageSpace"
      threshold = 5368709120 # 5 GiB in bytes
      operator  = "LessThanThreshold"
    }
    connections = {
      metric    = "DatabaseConnections"
      threshold = 200
      operator  = "GreaterThanThreshold"
    }
  }
}

resource "aws_cloudwatch_metric_alarm" "rds" {
  for_each = local.rds_alarms

  alarm_name          = "ditwl-pro-db-${each.key}"
  namespace           = "AWS/RDS"
  metric_name         = each.value.metric
  statistic           = "Average"
  period              = 300
  evaluation_periods  = 3
  threshold           = each.value.threshold
  comparison_operator = each.value.operator

  dimensions = {
    DBInstanceIdentifier = aws_db_instance.main.identifier
  }

  alarm_actions = [aws_sns_topic.alerts.arn]
}

This uses for_each over a map, so adding an alarm is adding an entry. See RDS with Terraform.

Turn log lines into metrics #

A metric filter counts occurrences of a pattern in a log group, and an alarm can use it:

metric-filter.tf
resource "aws_cloudwatch_log_metric_filter" "errors" {
  name           = "app-errors"
  log_group_name = aws_cloudwatch_log_group.app.name
  pattern        = "ERROR"

  metric_transformation {
    name      = "AppErrorCount"
    namespace = "Ditwl/App"
    value     = "1"
  }
}

resource "aws_cloudwatch_metric_alarm" "app_errors" {
  alarm_name          = "ditwl-pro-app-errors"
  namespace           = "Ditwl/App"
  metric_name         = "AppErrorCount"
  statistic           = "Sum"
  period              = 300
  evaluation_periods  = 1
  threshold           = 10
  comparison_operator = "GreaterThanOrEqualToThreshold"
  treat_missing_data  = "notBreaching"
  alarm_actions       = [aws_sns_topic.alerts.arn]
}

A dashboard #

dashboard.tf
resource "aws_cloudwatch_dashboard" "main" {
  dashboard_name = "ditwl-pro"

  dashboard_body = jsonencode({
    widgets = [
      {
        type   = "metric"
        x      = 0
        y      = 0
        width  = 12
        height = 6
        properties = {
          title  = "Web CPU"
          region = "eu-west-1"
          stat   = "Average"
          period = 300
          metrics = [
            ["AWS/EC2", "CPUUtilization", "InstanceId", aws_instance.web.id]
          ]
        }
      }
    ]
  })
}

What to alarm on #

Service Metrics
EC2 StatusCheckFailed, CPU, credit balance for burstable instances
RDS CPU, free storage, connections, replica lag
ALB HTTPCode_Target_5XX_Count, UnHealthyHostCount, latency
Lambda Errors, Throttles, Duration
SQS ApproximateAgeOfOldestMessage, DLQ messages visible
Billing EstimatedCharges (in us-east-1)

Alarm on symptoms that need action. Too many alarms are ignored. Costs: alarms, custom metrics, dashboards and logs ingestion are billed, so review them with Infracost.

#AWS #Cloudwatch #Terraform #OpenTofu