Terraform Provisioners and EC2 user_data: When to Use Each (and Alternatives)

· 2 min read · Terraform & OpenTofu Tutorials

Configure a server after creating it #

Terraform creates infrastructure. Sometimes you also need to install software or run a script on the new machine. There are four ways, from the most to the least recommended:

  1. An image that already contains everything (AMI built with Packer).
  2. user_data with cloud-init, executed at the first boot.
  3. A configuration tool such as Ansible.
  4. Provisioners, a last resort.

user_data #

user_data is a script or cloud-init file that the instance runs at first boot. It is part of the resource, so it is planned and tracked.

ec2.tf
resource "aws_instance" "web" {
  ami                    = data.aws_ami.ubuntu.id
  instance_type          = "t3.micro"
  subnet_id              = aws_subnet.private[0].id
  vpc_security_group_ids = [aws_security_group.web.id]

  user_data = templatefile("${path.module}/user-data.sh.tftpl", {
    app_version = var.app_version
  })

  user_data_replace_on_change = true

  tags = {
    Name = "ditwl-web"
  }
}
user-data.sh.tftpl
#!/bin/bash
set -euo pipefail
apt-get update -y
apt-get install -y nginx
echo "Version ${app_version}" > /var/www/html/index.html
systemctl enable --now nginx
  • templatefile injects Terraform values into the script.
  • user_data_replace_on_change = true recreates the instance when the script changes. Without it, a change would not apply because the script only runs on the first boot.
  • User data is limited to 16 KB and is visible to anyone who can describe the instance, so never put secrets in it. Read them at runtime from Secrets Manager.
  • Check the result in /var/log/cloud-init-output.log on the instance. See also EC2 with Terraform.

Provisioners #

Provisioners run commands during creation or destruction of a resource.

provisioner.tf
resource "aws_instance" "web" {
  ami           = data.aws_ami.ubuntu.id
  instance_type = "t3.micro"
  key_name      = aws_key_pair.deployer.key_name

  connection {
    type        = "ssh"
    user        = "ubuntu"
    private_key = file(var.private_key_path)
    host        = self.public_ip
  }

  provisioner "file" {
    source      = "app.conf"
    destination = "/tmp/app.conf"
  }

  provisioner "remote-exec" {
    inline = [
      "sudo apt-get update -y",
      "sudo apt-get install -y nginx",
      "sudo mv /tmp/app.conf /etc/nginx/conf.d/app.conf",
      "sudo systemctl restart nginx",
    ]
  }

  provisioner "local-exec" {
    command = "echo ${self.public_ip} >> inventory.txt"
  }
}
Provisioner What it does
local-exec Runs a command on the machine that runs Terraform
remote-exec Runs commands on the remote resource through SSH or WinRM
file Copies files to the remote resource

Why provisioners are a last resort #

  • They are not declarative: Terraform does not know what the script changed and cannot detect drift.
  • They need network access (SSH or WinRM) from the machine running Terraform to the instance, which usually means public IPs or a bastion.
  • If a provisioner fails, the resource is marked tainted and is recreated on the next apply.
  • They only run on creation (or destruction with when = destroy), so changes to the script do nothing to existing resources.
  • Errors are hard to debug.

HashiCorp's own documentation recommends them only when the provider cannot do what you need.

Reasonable uses #

  • A local-exec to trigger something outside Terraform, such as calling an API or running a local script after creation.
  • A local-exec with when = destroy for cleanup.
  • Use terraform_data with triggers_replace to control when a provisioner runs again:
terraform_data.tf
resource "terraform_data" "register" {
  triggers_replace = [aws_instance.web.id, var.app_version]

  provisioner "local-exec" {
    command = "./register.sh ${aws_instance.web.private_ip}"
  }
}

terraform_data replaces the older null_resource.

Better alternatives #

Related: lifecycle and replace and taint.

#Terraform #OpenTofu #AWS #Ansible