AWS with Terraform Tutorial: AWS Auto Scaling (15)
How to create AWS Auto Scaling Groups with Terraform #
Using the Terraform aws_launch_template, aws_autoscaling_group and aws_autoscaling_policy resource blocks to run a self-healing group of EC2 instances that grows and shrinks with demand.
Welcome to our tutorial series about Terraform or OpenTofu on AWS. In the previous sections a VPC, subnets, security groups, EC2 instances, a database and DNS were defined. In this section the front-end server stops being a single instance and becomes a group of instances that AWS replaces when they fail and scales when the load changes.
Prerequisites #
Read the previous sections of the tutorial, listed in the series index at the end of this page. This section reuses:
- the public subnets
ditwl-sn-za-pro-pub-00andditwl-sn-zb-pro-pub-04from AWS Subnets, - the security groups
ditwl-sg-base-ec2andditwl-sg-front-endfrom AWS Security Groups, - the key pair
ditwl-kp-config-userfrom AWS Key Pairs, - the AMI data source
ubuntu-23-04-arm64-minimalfrom AWS AMIs.
AWS Auto Scaling #
An Amazon EC2 Auto Scaling group (ASG) is a collection of EC2 instances managed as a single unit. It keeps the number of instances between a minimum and a maximum, spreads them across Availability Zones, replaces the ones that fail their health checks and adds or removes instances following scaling policies.
Three resources are needed:
- Launch template (
aws_launch_template): the blueprint of every instance in the group: AMI, instance type, key pair, security groups, user data and tags. - Auto Scaling group (
aws_autoscaling_group): how many instances, in which subnets (Availability Zones) and how to check their health. - Scaling policy (
aws_autoscaling_policy) and optionally scheduled actions (aws_autoscaling_schedule): when to add or remove instances.
What changes in the infrastructure #
The EC2 section defined one front-end server, ditwl-ec-front-end-001. Instances in an Auto Scaling group are created and destroyed by AWS, so they cannot be defined one by one in Terraform and their IP addresses change. Make these changes in terraform-aws-tutorial.tf:
- Remove (or comment out) the resource
aws_instance.ditwl-ec-front-end-001. The back-end serverditwl-ec-back-end-123is not changed. - Remove the Route 53 records that point to that instance (
ditwl-r53-public-front-end-001andditwl-r53-private-front-end-001). The next section, AWS Load Balancers, adds one stable DNS name for the whole group.
Definition of an Auto Scaling group with Terraform #
Launch template: ditwl-lt-front-end #
# Launch template for the front-end servers
resource "aws_launch_template" "ditwl-lt-front-end" {
name = "ditwl-lt-front-end"
image_id = data.aws_ami.ubuntu-23-04-arm64-minimal.id
instance_type = "t4g.micro"
key_name = "ditwl-kp-config-user"
vpc_security_group_ids = [aws_security_group.ditwl-sg-base-ec2.id, aws_security_group.ditwl-sg-front-end.id]
update_default_version = true # every change creates a new version and makes it the default
metadata_options {
http_endpoint = "enabled"
http_tokens = "required" # only IMDSv2 (session tokens) can read the instance metadata
}
# Cloud-init script: install a web server that shows the name of the instance
user_data = base64encode(<<-EOT
#!/bin/bash
apt-get update
apt-get install -y nginx
echo "front-end $(hostname)" > /var/www/html/index.html
EOT
)
tag_specifications {
resource_type = "instance"
tags = {
"Name" = "ditwl-ec-front-end"
"app" = "front-end"
"os" = "ubuntu"
"environment" = "pro"
"cost_center" = "marketing-department"
"owner" = "IT Wonder Lab"
}
}
}user_datamust be Base64 encoded in launch templates.base64encode()takes care of it.http_tokens = "required"is a good security default: it disables IMDSv1, which is used in server-side request forgery attacks.- The provider
default_tagsare not copied to the instances launched by an Auto Scaling group, that is why the tags are repeated intag_specifications.
Auto Scaling group: ditwl-asg-front-end #
# Auto Scaling group for the front-end servers, one or more instances in two Availability Zones
resource "aws_autoscaling_group" "ditwl-asg-front-end" {
name = "ditwl-asg-front-end"
min_size = 1
max_size = 4
desired_capacity = 2
vpc_zone_identifier = [aws_subnet.ditwl-sn-za-pro-pub-00.id, aws_subnet.ditwl-sn-zb-pro-pub-04.id]
health_check_type = "EC2"
health_check_grace_period = 120
launch_template {
id = aws_launch_template.ditwl-lt-front-end.id
version = aws_launch_template.ditwl-lt-front-end.latest_version
}
# Replace the instances gradually when the launch template changes (new AMI, new user data...)
instance_refresh {
strategy = "Rolling"
preferences {
min_healthy_percentage = 50
}
}
# Tag applied to the group and to every instance it launches
tag {
key = "Name"
value = "ditwl-ec-front-end"
propagate_at_launch = true
}
# The scaling policy changes the desired capacity, Terraform must not undo it
lifecycle {
ignore_changes = [desired_capacity]
}
}min_size,max_sizeanddesired_capacityare the lower limit, the upper limit (also a cost guardrail) and the number of instances at the start.vpc_zone_identifierlists one subnet per Availability Zone. The group balances the instances between them.version = ...latest_versionmakes the group change when the launch template changes, which is what triggers the instance refresh. With"$Latest"Terraform would not see any difference.
Scaling policy: keep the average CPU at 50% #
# Target tracking: AWS adds or removes instances to keep the average CPU around 50%
resource "aws_autoscaling_policy" "ditwl-asp-front-end-cpu" {
name = "ditwl-asp-front-end-cpu"
autoscaling_group_name = aws_autoscaling_group.ditwl-asg-front-end.name
policy_type = "TargetTrackingScaling"
estimated_instance_warmup = 120
target_tracking_configuration {
predefined_metric_specification {
predefined_metric_type = "ASGAverageCPUUtilization"
}
target_value = 50.0
}
}Target tracking works like a thermostat: it creates the CloudWatch alarms and calculates how many instances are needed. It is the simplest policy and the right choice in most cases.
Scheduled action: fewer instances at night #
# Every day at 22:00 UTC reduce the group to one instance (and back to two at 06:00 UTC)
resource "aws_autoscaling_schedule" "ditwl-ass-front-end-night" {
scheduled_action_name = "ditwl-ass-front-end-night"
autoscaling_group_name = aws_autoscaling_group.ditwl-asg-front-end.name
min_size = 1
max_size = 2
desired_capacity = 1
recurrence = "0 22 * * *"
}
resource "aws_autoscaling_schedule" "ditwl-ass-front-end-day" {
scheduled_action_name = "ditwl-ass-front-end-day"
autoscaling_group_name = aws_autoscaling_group.ditwl-asg-front-end.name
min_size = 1
max_size = 4
desired_capacity = 2
recurrence = "0 6 * * *"
}Run the Terraform Plan #
Plan using OpenTofu #
$ tofu plan
The plan creates the launch template, the Auto Scaling group, the scaling policy and the two scheduled actions, and destroys the instance ditwl-ec-front-end-001 (and the DNS records if they were defined).
Apply the changes using OpenTofu #
$ tofu apply
Type yes to confirm. The group starts two instances, one in each Availability Zone. List them with the AWS CLI:
$ aws autoscaling describe-auto-scaling-groups \
--auto-scaling-group-names ditwl-asg-front-end \
--query 'AutoScalingGroups[0].Instances[].[InstanceId,AvailabilityZone,LifecycleState,HealthStatus]' \
--output table --profile ditwl_infradmin
Test the self-healing #
Terminate one of the instances and watch the group create a replacement in a couple of minutes, without any change in Terraform:
$ aws ec2 terminate-instances --instance-ids i-0123456789abcdef0 --profile ditwl_infradmin
$ aws autoscaling describe-scaling-activities --auto-scaling-group-name ditwl-asg-front-end \
--max-items 3 --profile ditwl_infradmin
To see the scale-out, connect to one instance and generate CPU load (for example with stress-ng). After a few minutes the average CPU goes over 50% and the group adds instances up to max_size.
Destroy the infrastructure using OpenTofu #
$ tofu destroy
AWS Auto Scaling Cost #
An Auto Scaling group has no additional charge: you pay for the resources it launches (EC2 instances, EBS volumes and data transfer). The cost is therefore controlled by max_size and the instance type. Detailed CloudWatch monitoring, if enabled in the launch template, is billed separately. Check the EC2 pricing page for the current prices and remember to destroy the infrastructure after the tests.
Common Questions About AWS Auto Scaling #
What is the difference between minimum, maximum and desired capacity? #
min_size and max_size are the limits that no policy can cross. desired_capacity is the number of instances the group tries to have right now. Scaling policies change the desired capacity inside those limits.
Why is desired_capacity in ignore_changes? #
Because AWS changes it while scaling. Without ignore_changes, the next tofu apply would reset the group to the value in the code and fight with the scaling policy.
How can I roll out a new AMI or a new version of the user data? #
Change the launch template. The new version is used by new instances, and the instance_refresh block replaces the existing ones gradually, keeping at least 50% of the capacity healthy.
Can I use Spot Instances to reduce the cost? #
Yes. Replace the launch_template block with a mixed_instances_policy block, which lets the group combine On-Demand and Spot capacity and several instance types. It is a good fit for stateless front-end servers.
Why not use count or for_each on aws_instance? #
Terraform would create a fixed set of instances and nothing would replace a failed one or react to the load. An Auto Scaling group delegates that work to AWS.
Next Steps #
The instances in the group are created and destroyed all the time and each one has a different IP address, so users need a single stable entry point. Continue with AWS Load Balancers.