Infrastructure as Code (IaC) means describing servers, networks, databases, permissions and DNS in text files that live in version control, and letting a tool make reality match those files. The promise is that infrastructure gets the same treatment as application code: review, history, repeatability, rollback. The reality is more specific. IaC tools are diff engines with a memory, and nearly every IaC incident, from a deleted database to a week of drift, comes from misunderstanding what the diff compares and where the memory lives.

This article explains the model from first principles, works through a real Terraform configuration, and covers state, refactoring, pipeline design, testing and the failure modes worth designing against. The examples use Terraform syntax, which OpenTofu shares; the concepts carry over to Pulumi, CloudFormation and Bicep.

Advertisement

Declarative desired state versus scripts

The pre-IaC approach is an imperative script: create this VPC, then this subnet, then this instance. It works once. Run it again and it fails, or creates duplicates, because it describes actions, not outcomes. Making it safe to rerun means writing existence checks for every resource, and making it handle changes means writing update logic for every attribute. You end up writing a bad IaC tool.

A declarative tool inverts this. You describe the end state: three buckets with these names, versioning on two of them. The tool computes the actions needed to get from what exists to what you described. Running it twice is safe because the second run finds nothing to do, which is the property called idempotence. Changing an attribute in the file produces an update; deleting a block produces a destroy. The file is both the documentation and the mechanism.

Configuration management tools such as Ansible sit in between: their modules are mostly idempotent, but a playbook is still an ordered list of tasks, and removing a task does not remove what it created. That difference is why provisioning and configuration are usually handled by different tools.

The architecture: three views and a diff

A declarative IaC tool juggles three views of the world: the configuration you wrote, the state the tool recorded last time, and the real infrastructure as reported by cloud APIs. State exists because the configuration alone cannot tell the tool which real resource corresponds to which block. The bucket acme-raw might be yours, or someone else's with the same name. State is the mapping from resource addresses like aws_s3_bucket.data["raw"] to real IDs, plus the last-known attributes.

The desired-state loop: three views of the world, one diffConfiguration (git)what you wantState filewhat the tool last sawReal infrastructurewhat exists in the cloudparse + graphreadrefresh via APIsPlancreate / update / replace / destroyreviewed, saved planApplywalk graph, call provider APIsAPI callsState updatedunder a lockPlan compares all three: configuration against state, state refreshed against reality.Apply only executes the plan; anything not in the configuration is not the tool's business.Lose the state and the tool forgets it owns resources; edit reality by hand and the next plan reverts it.
Plan refreshes state against reality, diffs it against the configuration, and produces a change set. Apply executes that change set in dependency order and records the result.

From the configuration the tool builds a dependency graph. Every reference, such as one resource reading another's id, becomes an edge. Apply walks that graph, running independent nodes in parallel and dependent ones in order, and destroys in reverse order. You rarely state ordering explicitly; you get it by referencing values. When ordering exists that no reference expresses, such as an IAM policy that must exist before a service can use it, you add an explicit depends_on.

Advertisement

Worked example: three buckets

Here is a complete configuration: a pinned provider, a remote backend with locking, and buckets generated from a map.

terraform {
  required_version = ">= 1.10"
  required_providers {
    aws = { source = "hashicorp/aws", version = "~> 5.0" }
  }
  backend "s3" {
    bucket       = "acme-tfstate-prod"
    key          = "data/buckets.tfstate"
    region       = "eu-west-1"
    encrypt      = true
    use_lockfile = true          # S3-native lock (1.10+); replaces the deprecated dynamodb_table
  }
}

variable "buckets" {
  type = map(object({ versioned = bool }))
  default = {
    raw     = { versioned = true }
    curated = { versioned = true }
    scratch = { versioned = false }
  }
}

resource "aws_s3_bucket" "data" {
  for_each = var.buckets              # keyed by name, not by list position
  bucket   = "acme-${each.key}"
  lifecycle { prevent_destroy = true }
}

resource "aws_s3_bucket_versioning" "data" {
  for_each = { for k, v in var.buckets : k => v if v.versioned }
  bucket   = aws_s3_bucket.data[each.key].id     # this reference is a graph edge
  versioning_configuration { status = "Enabled" }
}

On the first terraform plan the state is empty, so the plan shows five creates: three buckets and two versioning resources. Apply runs the buckets in parallel, then the versioning resources, because each references its bucket's ID. Now change scratch to versioned = true. The next plan shows exactly one create. Delete the scratch entry from the map and the plan shows a destroy, which prevent_destroy turns into an error, a deliberate speed bump on stateful resources.

The plan symbols are the most important thing to read in any change: + create, ~ update in place, - destroy, and -/+ replace, meaning destroy and recreate. Some attributes cannot be changed in place, a bucket's name or a database's engine for instance, so changing them forces replacement. A one-line diff that shows -/+ on a database is a data-loss event waiting for an approval click.

Why for_each over a map rather than count over a list? With count, resources are addressed by position: [0], [1], [2]. Remove the first list item and every later item shifts index, so the plan proposes destroying and recreating them. Keys make addresses stable.

State: the memory you must protect

State is the tool's only memory of ownership, so treat it like a database. Keep it in a remote backend, never in git or on a laptop, so a team shares one copy. Turn on locking, so two applies cannot interleave and corrupt it; the example uses S3's native lock file, available from Terraform 1.10, rather than the older DynamoDB table approach, which is now deprecated. Enable versioning on the state bucket, so a bad write can be rolled back. Encrypt it and restrict access, because state contains every attribute of every resource, including generated passwords and keys in plain text.

Never edit state by hand. Use the tool's state commands or, better, the declarative blocks in the next section, which go through plan and review. If an apply is killed mid-run, the lock may remain; terraform force-unlock clears it, but only after confirming nobody else is running.

Split state by blast radius. One giant state for a whole company makes every plan slow and every mistake global. A common split is per environment and per layer: network, data, compute and applications, each with its own state, exchanging values through outputs.

Refactoring without destroying anything

Code gets renamed and reorganised. In IaC a rename changes the resource address, and to a naive diff that looks like destroying the old resource and creating a new one. Three declarative blocks fix this, and because they are part of the configuration they appear in the plan and get reviewed.

# Renamed a resource: move its state instead of destroying and recreating it (Terraform 1.1+).
moved {
  from = aws_s3_bucket.raw_bucket
  to   = aws_s3_bucket.data["raw"]
}

# Adopt a bucket someone created by hand (Terraform 1.5+). Plan shows an import, not a create.
import {
  to = aws_s3_bucket.data["curated"]
  id = "acme-curated"
}

# Stop managing a resource without deleting it (Terraform 1.7+).
removed {
  from = aws_s3_bucket.legacy_logs
  lifecycle { destroy = false }
}

Use moved for every rename or move into a module, import to adopt hand-built resources, and removed to hand a resource to another state or team. In each case the plan should show no creates or destroys for the affected resource. If it does, stop.

Structuring code: modules and environments

A module is a directory of configuration with inputs and outputs, the IaC equivalent of a function. Good modules encode a decision, such as an encrypted bucket with logging and a deny-public policy, so teams cannot forget it. Bad modules are thin wrappers that pass through every argument and add only indirection. Version shared modules and pin the version, so a module change rolls out one environment at a time.

For environments, the most robust pattern is one directory per environment, each with its own backend and a small file of values, all calling the same modules. Workspaces, which give one directory several states, are fine for short-lived copies but make it too easy to apply production values while thinking you are in staging. Commit the dependency lock file, .terraform.lock.hcl, so every machine uses the same provider builds, and upgrade providers deliberately in their own change.

The pipeline: plan in review, apply what was reviewed

The core rule of an IaC pipeline is that the change applied must be exactly the change reviewed. So plan on every pull request, show the plan in the review, save it to a file, and apply that file after approval. If anything changed in between, the saved plan is rejected as stale rather than silently doing something new. Applies run from the pipeline only, using a role humans do not hold day to day. This is the same discipline as in a general CI/CD pipeline, with a higher cost of error.

#!/usr/bin/env bash
set -euo pipefail
terraform fmt -check -recursive
terraform init -input=false
terraform validate

# Exit codes with -detailed-exitcode: 0 = no changes, 1 = error, 2 = changes present.
set +e
terraform plan -input=false -lock-timeout=5m -out=tfplan -detailed-exitcode
rc=$?
set -e
[ "$rc" -eq 1 ] && exit 1

terraform show -json tfplan > plan.json
conftest test plan.json            # policy as code: e.g. deny any delete of a stateful resource

# Later, in a separate job after human approval, apply EXACTLY the reviewed plan file:
#   terraform apply -input=false tfplan

Policy as code turns review comments into automated rules: deny public buckets, require tags, refuse any plan that destroys a database. Tools such as Open Policy Agent with conftest evaluate the JSON plan. The same -detailed-exitcode run, scheduled nightly against production, is a cheap drift detector: exit code 2 means reality and code disagree. Deciding what to do about it is the subject of config drift reconciliation.

Testing infrastructure code

Testing works in layers. Static checks, formatting, validation and linters, catch typos and invalid arguments in seconds. Policy tests on the plan catch dangerous changes without creating anything. Module tests, written with terraform test and .tftest.hcl files or with tools such as Terratest, create real resources in a sandbox account, assert on them, and destroy them. They are slow and cost money, so reserve them for shared modules. Finally, a staging environment built from the same modules is the integration test. Applying to staging first and production second is the most valuable test you have.

Failure modes

FailureWhat happensDefence
Forced replacementAn innocent edit replaces a database or bucketRead -/+ in every plan; prevent_destroy; policy rules on destroys
Index shiftcount list reordered, resources recreatedUse for_each with stable keys
Lost or corrupted stateTool forgets ownership, tries to recreate everythingRemote, versioned, encrypted, locked backend; backups
Partial applyRun fails halfway, some resources changedRe-plan and apply; design changes to be safely re-runnable
Click-ops driftManual fix overwritten on next applyRead-only console for humans; nightly drift plan
Provider upgrade diffNew provider version proposes unexpected changesPin versions; upgrade alone in its own reviewed change
Secrets in state or codeCredentials readable by anyone with state accessSecret managers, generated secrets, restricted state access
Eventual consistencyApply fails because an API has not yet seen a new resourceRetry; explicit dependencies; avoid tight create-then-use loops

Choosing a tool

ToolModelBest fit
Terraform / OpenTofuHCL, explicit state, many providersMulti-cloud and SaaS; the broadest ecosystem. OpenTofu is the open-source fork created after HashiCorp's 2023 licence change
PulumiGeneral-purpose languages, explicit stateTeams wanting loops, types and tests in a familiar language
CloudFormation / AWS CDKService-managed stateAWS-only estates; no state bucket to run
Bicep / ARMService-managed deploymentsAzure-only estates
CrossplaneKubernetes controllers reconcile continuouslyPlatform teams already running everything through Kubernetes

The deepest trade-off is between plan-and-apply tools, which change things only when you run them, and continuous reconcilers, which fix drift automatically but can also fight humans during an incident. Neither is wrong. Choose deliberately, and make sure people know which model they are in.

What to do next

  1. Move any local state to a remote, versioned, encrypted backend with locking enabled.
  2. Make the pipeline plan on every pull request, show the plan, and apply only the saved, approved plan file.
  3. Add policy rules that block destroys and replacements of stateful resources without an explicit override.
  4. Replace count over lists with for_each over maps, using moved blocks so nothing is recreated.
  5. Schedule a nightly -detailed-exitcode plan against production and alert on exit code 2.
  6. Pin provider and module versions, commit the lock file, and upgrade them in dedicated changes.
Key takeaway: Infrastructure as Code is a diff engine with a memory: it compares your configuration with recorded state refreshed against real infrastructure, then executes the difference over a dependency graph. Safety comes from protecting the memory with a remote, locked, versioned state backend; from reading every plan, especially replacements; from refactoring with moved, import and removed blocks instead of deletes; and from a pipeline that applies exactly the reviewed plan, gated by policy and followed by drift checks.