AWS CloudWatch Alarms, Logs and Dashboards with Terraform
Observability as code #
Amazon CloudWatch collects metrics, logs and events from AWS. Defining alarms and dashboards in Terraform means that every new resource ships with its monitoring and that it is reviewed like any other change.
An SNS topic to notify #
Alarms need somewhere to send the notification. Create a topic and subscribe the operations team (see SQS and SNS):
resource "aws_sns_topic" "alerts" {
name = "ditwl-pro-alerts"
}
resource "aws_sns_topic_subscription" "ops" {
topic_arn = aws_sns_topic.alerts.arn
protocol = "email"
endpoint = "ops@example.com"
}Confirm the subscription with the link that arrives by email.
Log groups with retention #
By default, log groups keep logs forever and cost grows silently. Always set a retention:
resource "aws_cloudwatch_log_group" "app" {
name = "/ditwl/pro/app"
retention_in_days = 30
kms_key_id = aws_kms_key.logs.arn # optional
}An alarm for high CPU on EC2 #
resource "aws_cloudwatch_metric_alarm" "ec2_cpu" {
alarm_name = "ditwl-pro-web-cpu-high"
alarm_description = "CPU above 80% for 10 minutes"
namespace = "AWS/EC2"
metric_name = "CPUUtilization"
statistic = "Average"
period = 300
evaluation_periods = 2
threshold = 80
comparison_operator = "GreaterThanThreshold"
treat_missing_data = "notBreaching"
dimensions = {
InstanceId = aws_instance.web.id
}
alarm_actions = [aws_sns_topic.alerts.arn]
ok_actions = [aws_sns_topic.alerts.arn]
}period is the length of each data point in seconds, and the alarm goes to ALARM when evaluation_periods consecutive points are above the threshold. treat_missing_data decides what happens when there is no data, which is important for stopped instances.
Alarms for several resources with for_each #
locals {
rds_alarms = {
cpu = {
metric = "CPUUtilization"
threshold = 80
operator = "GreaterThanThreshold"
}
free_storage = {
metric = "FreeStorageSpace"
threshold = 5368709120 # 5 GiB in bytes
operator = "LessThanThreshold"
}
connections = {
metric = "DatabaseConnections"
threshold = 200
operator = "GreaterThanThreshold"
}
}
}
resource "aws_cloudwatch_metric_alarm" "rds" {
for_each = local.rds_alarms
alarm_name = "ditwl-pro-db-${each.key}"
namespace = "AWS/RDS"
metric_name = each.value.metric
statistic = "Average"
period = 300
evaluation_periods = 3
threshold = each.value.threshold
comparison_operator = each.value.operator
dimensions = {
DBInstanceIdentifier = aws_db_instance.main.identifier
}
alarm_actions = [aws_sns_topic.alerts.arn]
}This uses for_each over a map, so adding an alarm is adding an entry. See RDS with Terraform.
Turn log lines into metrics #
A metric filter counts occurrences of a pattern in a log group, and an alarm can use it:
resource "aws_cloudwatch_log_metric_filter" "errors" {
name = "app-errors"
log_group_name = aws_cloudwatch_log_group.app.name
pattern = "ERROR"
metric_transformation {
name = "AppErrorCount"
namespace = "Ditwl/App"
value = "1"
}
}
resource "aws_cloudwatch_metric_alarm" "app_errors" {
alarm_name = "ditwl-pro-app-errors"
namespace = "Ditwl/App"
metric_name = "AppErrorCount"
statistic = "Sum"
period = 300
evaluation_periods = 1
threshold = 10
comparison_operator = "GreaterThanOrEqualToThreshold"
treat_missing_data = "notBreaching"
alarm_actions = [aws_sns_topic.alerts.arn]
}A dashboard #
resource "aws_cloudwatch_dashboard" "main" {
dashboard_name = "ditwl-pro"
dashboard_body = jsonencode({
widgets = [
{
type = "metric"
x = 0
y = 0
width = 12
height = 6
properties = {
title = "Web CPU"
region = "eu-west-1"
stat = "Average"
period = 300
metrics = [
["AWS/EC2", "CPUUtilization", "InstanceId", aws_instance.web.id]
]
}
}
]
})
}What to alarm on #
| Service | Metrics |
|---|---|
| EC2 | StatusCheckFailed, CPU, credit balance for burstable instances |
| RDS | CPU, free storage, connections, replica lag |
| ALB | HTTPCode_Target_5XX_Count, UnHealthyHostCount, latency |
| Lambda | Errors, Throttles, Duration |
| SQS | ApproximateAgeOfOldestMessage, DLQ messages visible |
| Billing | EstimatedCharges (in us-east-1) |
Alarm on symptoms that need action. Too many alarms are ignored. Costs: alarms, custom metrics, dashboards and logs ingestion are billed, so review them with Infracost.