Terraform lets you describe infrastructure as text, then makes the real world match that text. You do not write the steps to create a network, a bucket and a database in order; you write down that they should exist with certain properties, and Terraform works out what to create, change or delete, in what order, by calling cloud provider APIs. The same files create a fresh environment, update an existing one and tell you in advance exactly what an update will do.

That last property is the reason teams adopt it. A reviewed plan turns an infrastructure change from a sequence of console clicks into a diff that can be read, approved and repeated. This article builds the mental model from first principles, walks a worked example, and then spends most of its time on the parts that cause real incidents: state, locking, replacement, drift and refactoring. If you are new to the wider practice, the DevOps introduction sets the context.

Advertisement

Three things that must agree

Terraform reasons about three separate pictures of your infrastructure. The configuration is the set of .tf files you write: the desired state. The real infrastructure is whatever the provider APIs report right now. The state file sits between them: a JSON record of every object Terraform created, mapping each resource address in your code, such as aws_s3_bucket.artifacts["ml"], to a real-world identifier such as a bucket name or an instance ID, plus the attributes last seen.

State is what makes Terraform more than a script. Without it, Terraform could not know that the bucket named in your code is the one it created last week rather than an unrelated one, could not detect that you deleted a resource from the code and should therefore delete it from the cloud, and could not compute a diff cheaply. Every important behaviour, and most of the incidents, trace back to the relationship between these three pictures.

The Terraform loop: desired config, recorded state and real infrastructure are reconciled by a planConfiguration (.tf)what you wantState filewhat Terraform createdProvider APIswhat actually existsterraform planbuild graph, refresh, diffrefreshSaved plancreate / update / replace / destroyterraform applywalk graph, parallel API callsreview, approveNew statewritten back, lock releasednext run reads itThe lockremote backend holds a lock during plan and apply, so two runs never write the same state at onceDrift is the gap between the state box and the provider box; refresh measures it, the next apply closes it.
Each run reads configuration and state, refreshes state against provider APIs, builds a dependency graph and produces a plan; apply executes the plan and writes the new state, all under a lock.

How a run works

A run has four stages. First, terraform init downloads the provider plugins named in the configuration (the AWS provider, the Kubernetes provider and so on) and configures the backend where state lives. Providers are separate binaries that translate Terraform resources into API calls; Terraform core knows nothing about buckets or VMs.

Second, plan. Terraform parses the configuration and builds a directed acyclic graph of resources. An edge exists whenever one resource refers to another: the versioning resource below reads each.value.id from the bucket, so the bucket must exist first. It then refreshes, asking providers for the current attributes of everything in state, and diffs desired against current. Each resource gets an action: create, update in place, destroy, or replace (destroy and create) when an attribute cannot be changed in place.

Third, apply walks the graph, running independent branches concurrently (ten at a time by default, adjustable with -parallelism), and records each result in state as it goes. Fourth, the new state is persisted and the lock released. Because the graph must be acyclic, a reference loop between two resources is a configuration error that Terraform reports before calling any API; the general technique is the same one covered in topological sorting.

Advertisement

A worked example

Suppose each team needs a versioned artifact bucket per environment. The configuration below declares the provider, a remote backend, three input variables, a bucket per team with for_each, versioning for each bucket, and an output map of ARNs.

terraform {
  required_version = ">= 1.11"
  required_providers {
    aws = { source = "hashicorp/aws", version = "~> 6.0" }
  }
  backend "s3" {
    bucket       = "acme-tfstate"
    key          = "artifacts/prod/terraform.tfstate"
    region       = "eu-west-1"
    encrypt      = true
    use_lockfile = true      # S3-native lock (.tflock object)
  }
}

provider "aws" { region = var.region }

variable "region"      { type = string }
variable "environment" { type = string }
variable "teams"       { type = set(string) }

resource "aws_s3_bucket" "artifacts" {
  for_each = var.teams
  bucket   = "acme-${var.environment}-${each.key}-artifacts"
  tags     = { team = each.key, managed_by = "terraform" }
}

resource "aws_s3_bucket_versioning" "artifacts" {
  for_each = aws_s3_bucket.artifacts
  bucket   = each.value.id
  versioning_configuration { status = "Enabled" }
}

output "bucket_arns" {
  value = { for k, b in aws_s3_bucket.artifacts : k => b.arn }
}

With teams = ["ml", "web"] the first plan proposes four creates: two buckets and two versioning resources. Terraform knows the versioning resources depend on the buckets and orders them accordingly. Add "data" to the set and the next plan shows exactly two creates for the new team and nothing else. Remove "web" and the plan shows two destroys. That is the whole value proposition in miniature: change the description, read the diff, apply.

Notice for_each over a set rather than count over a list. With count, resources are addressed by position, so deleting the first element of a list shifts every later element down one slot, and Terraform plans to destroy and recreate buckets whose names did not change. With for_each, addresses are keyed by name and stay stable. Use count only for an on/off toggle (zero or one).

The daily workflow

Teams converge on the same five commands. The important habit is saving the plan to a file and applying that file, so what gets applied is byte-for-byte what was reviewed, not a fresh plan computed minutes later against a world that may have changed.

terraform init                       # download providers, configure backend
terraform fmt -check && terraform validate
terraform plan -var-file=prod.tfvars -out=tfplan
terraform show tfplan                # what a reviewer reads
terraform apply tfplan               # applies exactly the reviewed plan

In CI, run fmt, validate and plan on every pull request and post the plan output as a comment; run apply only on merge, from the saved plan artifact, with credentials that only the pipeline holds. The CI/CD guide covers pipeline structure in general; the Terraform-specific rule is that humans review plans and machines apply them.

State, backends and locking

By default state is a local file, which is fine for experiments and wrong for teams: two people with two copies of state will each believe they own the infrastructure. A remote backend stores state centrally, in S3, Azure Blob Storage, Google Cloud Storage or HCP Terraform, and serialises writers with a lock. If a second run starts while the first holds the lock, it fails fast instead of interleaving writes.

For the S3 backend, Terraform 1.10 introduced S3-native locking as an experiment and 1.11 made it generally available: with use_lockfile = true the lock is a .tflock object written next to the state using S3 conditional writes, and HashiCorp has deprecated the older DynamoDB-table locking in its favour. If your configuration still names a dynamodb_table, both locks can run together while you migrate.

Three rules about state are non-negotiable. It contains secrets in plain text, because providers record attributes such as generated database passwords, so the bucket must be encrypted, versioned and readable only by the pipeline. It should be split: one state per environment and per blast radius (network, data, application), so a mistake in one plan cannot touch everything. And it should never be edited by hand; use terraform state mv, rm and the refactoring blocks below, and keep bucket versioning on so a bad write can be rolled back.

Modules: the unit of reuse

A module is a directory of .tf files with variables as inputs and outputs as return values. The root module is the directory you run Terraform in; it calls child modules with a module block and a source that can be a local path, a Git URL or a registry address with a version constraint.

Good modules are small and opinionated: an artifact-bucket module that always enables versioning, encryption and public-access blocking is more valuable than a thin wrapper exposing every bucket argument, because it encodes a decision. Pin module and provider versions, keep the lock file .terraform.lock.hcl in version control so every machine resolves identical provider builds, and upgrade deliberately in a dedicated change whose plan you read carefully.

Refactoring without destroying anything

Code changes that look cosmetic can be destructive. Renaming a resource or moving it into a module changes its address, and Terraform's default reading is that the old address should be destroyed and a new one created. For a bucket full of data or a production database, that is an outage. Three blocks let you say what you actually mean:

# rename without destroy/create
moved {
  from = aws_s3_bucket.artifacts["ml"]
  to   = aws_s3_bucket.artifacts["ml-platform"]
}

# adopt a bucket someone created by hand
import {
  to = aws_s3_bucket.artifacts["legacy"]
  id = "acme-prod-legacy-artifacts"
}

# stop managing it, but leave it running
removed {
  from = aws_s3_bucket.artifacts["sandbox"]
  lifecycle { destroy = false }
}

A moved block tells Terraform the object at the old address is now at the new one; the plan shows a move and no destroy. An import block adopts an existing object into state, and terraform plan -generate-config-out=generated.tf can draft its configuration for you to clean up. A removed block with destroy = false drops an object from management without deleting it. All three are reviewed in the plan like any other change, which is far safer than the older imperative state mv commands.

Failure modes and how to read them

Most Terraform incidents fall into a handful of patterns, and nearly all of them are visible in the plan if someone reads it.

SymptomCausePrevention
Plan shows replace on a databaseAn immutable attribute changed (engine, subnet group, name)Read every -/+ line; add lifecycle prevent_destroy on stateful resources
Mass destroy/create after a list editcount indexed by positionUse for_each with stable keys; moved blocks when converting
Error acquiring the state lockConcurrent run, or a crashed run left the lockConfirm nobody is running, then terraform force-unlock with the lock ID
Plan shows changes nobody madeDrift: someone edited resources in the consoleScheduled plan -detailed-exitcode in CI; restrict console write access
Plan flaps on every runProvider normalises a value differently from your codeMatch the canonical form, or lifecycle ignore_changes for that attribute
Secrets found in stateProviders record sensitive attributesEncrypt and restrict the backend; prefer managed secret references

Two lifecycle settings deserve a word. prevent_destroy = true makes any plan that would destroy the resource fail, which is a cheap guard on databases and state buckets. create_before_destroy = true reverses the replacement order so the new object exists before the old one goes, which avoids downtime but requires names that can coexist.

Terraform among its neighbours

Terraform is declarative, multi-provider and state-based. CloudFormation and Bicep are declarative too but bound to one cloud, with state held by the platform; the Bicep and ARM guide shows that model. The CDK lets you write infrastructure in a general-purpose language that synthesises templates (AWS CDK in depth). Configuration-management tools such as Ansible configure software inside machines rather than provisioning the machines themselves (Ansible introduction); a common split is Terraform for the cluster and network, then Kubernetes manifests or Ansible for what runs on top.

Licensing matters for some organisations. In August 2023 HashiCorp moved Terraform to the Business Source License, effective from version 1.6, and OpenTofu is a community fork that stays under the Mozilla Public License, hosted by the Linux Foundation. Both read the same HCL for most configurations, but they now evolve separately, so pick one per codebase and test before switching.

What to do next

  1. Install Terraform, write the bucket example against a sandbox account with a single team, and run init, plan and apply; then run destroy and read that plan too.
  2. Move state to a remote backend with encryption, versioning and use_lockfile = true; start two applies at once and watch the second fail on the lock.
  3. Change a list-based count resource to for_each using moved blocks, and confirm the plan shows moves and zero destroys.
  4. Add prevent_destroy to every stateful resource you manage and commit .terraform.lock.hcl.
  5. Wire fmt, validate and plan into pull requests and apply-from-saved-plan into merges; add a nightly drift plan with -detailed-exitcode.
  6. Split one large state into per-environment, per-layer states before it grows past what one person can review in a single plan.
Key takeaway: Terraform reconciles three pictures of your infrastructure, the configuration you write, the state it recorded and what the provider APIs report, by building a dependency graph and producing a plan you can read before anything changes. Keep state remote, encrypted, locked and split; use for_each with stable keys; refactor with moved, import and removed blocks; and treat every replace line in a plan as a question to answer before you approve it.