Terraform Data Sources and terraform_remote_state with AWS Examples

· 2 min read · Terraform & OpenTofu Tutorials

Read information that Terraform does not manage #

A resource creates and manages an object. A data source only reads information about an object that already exists, or that is managed somewhere else, so you can use it in your configuration. Data sources are read on every plan.

Data source examples #

Find the latest Ubuntu AMI (see also the AWS AMI tutorial):

data.tf
data "aws_ami" "ubuntu" {
  most_recent = true
  owners      = ["099720109477"] # Canonical

  filter {
    name   = "name"
    values = ["ubuntu/images/hvm-ssd-gp3/ubuntu-noble-24.04-amd64-server-*"]
  }
}

data "aws_caller_identity" "current" {}

data "aws_region" "current" {}

data "aws_availability_zones" "available" {
  state = "available"
}

data "aws_vpc" "default" {
  default = true
}

Use them with the data. prefix:

main.tf
resource "aws_instance" "web" {
  ami               = data.aws_ami.ubuntu.id
  instance_type     = "t3.micro"
  availability_zone = data.aws_availability_zones.available.names[0]

  tags = {
    Account = data.aws_caller_identity.current.account_id
  }
}

Other useful ones: aws_iam_policy_document (build an IAM policy in HCL), aws_route53_zone, aws_secretsmanager_secret_version, aws_ssm_parameter and aws_eks_cluster.

Sharing values between configurations #

In a real project the infrastructure is split in several configurations (network, database, applications), each one with its own state. The application needs the VPC ID created by the network configuration. There are two common ways.

terraform_remote_state #

Reads the outputs of another configuration from its backend:

remote-state.tf
data "terraform_remote_state" "network" {
  backend = "s3"

  config = {
    bucket = "ditwl-terraform-state"
    key    = "pro/network/terraform.tfstate"
    region = "eu-west-1"
  }
}

resource "aws_instance" "app" {
  ami           = data.aws_ami.ubuntu.id
  instance_type = "t3.micro"
  subnet_id     = data.terraform_remote_state.network.outputs.private_subnet_ids[0]
}

The other configuration must declare output "private_subnet_ids". The drawbacks: the reader needs read access to the whole state, which can contain secrets, and the two configurations are tightly coupled.

SSM Parameter Store as the interface #

The producer publishes the values and the consumers read them with a data source:

producer.tf
resource "aws_ssm_parameter" "vpc_id" {
  name  = "/pro/network/vpc_id"
  type  = "String"
  value = aws_vpc.main.id
}
consumer.tf
data "aws_ssm_parameter" "vpc_id" {
  name = "/pro/network/vpc_id"
}

# data.aws_ssm_parameter.vpc_id.value

This limits access to just the published values and does not depend on the backend used by the other team. A third option is a data source that finds the resource by tags, such as data "aws_vpc" with a filter on tag:Name.

Notes #

  • A data source that cannot find exactly one result fails the plan. That is a good thing: it avoids silently using the wrong object.
  • If a data source depends on a resource that has not been created yet, its read is deferred to the apply and its values are "known after apply".
  • Run the configurations in order: the network before the applications.

Related: variables, outputs and locals and project structure.

#Terraform #OpenTofu #AWS