AWS CloudFormation turns a declarative template into real infrastructure and keeps track of it as a unit called a stack. Tools such as the AWS CDK and AWS SAM generate CloudFormation templates, so even teams that never write YAML depend on its engine: how it orders work, decides whether an update replaces a resource, rolls back on failure and handles resources it did not create. Teams that understand the engine ship changes calmly; teams that do not eventually delete a production database with a one-line rename.
This article explains the engine from first principles: template anatomy, the dependency graph, update behaviours, change sets, rollback and its stuck states, policies that protect data, signals, custom resources, cross-stack wiring, drift and StackSets. It includes a guarded deployment script, a worked example and a checklist. Quotas were checked against the AWS documentation on 2026-10-03.
Architecture at a glance
Template anatomy and quotas
A template has one required section, Resources, and several optional ones: Parameters for inputs, Mappings for lookup tables, Conditions for environment switches, Outputs for values exported to people or other stacks, Transform for macros such as SAM, and Metadata. Each resource has a logical ID, a type such as AWS::DynamoDB::Table, properties, and optional attributes: DependsOn, DeletionPolicy, UpdateReplacePolicy, CreationPolicy, UpdatePolicy and Condition.
The logical ID is the stack's handle on the resource; the physical ID is the real name or ARN in AWS. CloudFormation maps one to the other and stores that mapping in the stack. Rename a logical ID and, to CloudFormation, the old resource disappears and a new one appears: it creates the new one and deletes the old one. That single fact explains many incidents, especially with CDK, where refactoring a construct tree changes logical IDs.
Templates have hard quotas. A template can declare at most 500 resources, 200 parameters, 200 outputs and 200 mappings. A template body passed inline is limited to 51,200 bytes, and one referenced from S3 to 1 MB. A stack template can contain 60 dynamic references. An account can hold 2,000 stacks per Region by default. Large systems therefore split into several stacks or nested stacks long before they hit the resource limit, mostly for blast radius rather than quotas.
The dependency graph
CloudFormation does not run resources top to bottom. It builds a directed acyclic graph from references: !Ref and !GetAtt create implicit dependencies, DependsOn adds explicit ones. Independent resources are created in parallel, and each waits only for its inputs. Deletion walks the graph in reverse. Add DependsOn only for dependencies the graph cannot see, such as a route to an internet gateway that must exist before instances try to download packages. Circular references fail validation, which commonly happens with security groups that reference each other; break the cycle with separate AWS::EC2::SecurityGroupIngress resources.
Update behaviours and replacement
On update, CloudFormation compares the new template and parameters with the stored ones, and for each changed property applies the update behaviour documented for that property: No interruption, Some interruption (for example a reboot), or Replacement. Replacement means creating a new physical resource, repointing everything that references it, and deleting the old one during the cleanup phase. Changing a DynamoDB table's key schema, an RDS instance's identifier or an EC2 instance's subnet requires replacement.
Replacement interacts badly with custom names. If a template sets TableName: orders and a change requires replacement, CloudFormation cannot create a second table named orders while the first exists, and the update fails with an error saying it cannot update a stack when a custom-named resource requires replacing. Letting CloudFormation generate names avoids the clash, and passing generated names around through references and outputs is the idiomatic approach.
Change sets and automated guards
A change set is a preview: CloudFormation computes what would be added, modified, removed or replaced, without touching anything. Each modification lists the action and a Replacement field of True, False or Conditional. Conditional usually means the value depends on something only known at execution time. Review change sets as code review for infrastructure, and automate the review for dangerous cases: a pipeline should refuse to execute a change set that replaces or removes stateful resources unless a human approves.
import sys, boto3
STATEFUL = {"AWS::DynamoDB::Table", "AWS::RDS::DBInstance", "AWS::RDS::DBCluster",
"AWS::S3::Bucket", "AWS::EFS::FileSystem", "AWS::Kinesis::Stream"}
cfn = boto3.client("cloudformation")
def deploy(stack, template_url, params, approved=False):
name = "ci-" + params["BuildId"]
cfn.create_change_set(StackName=stack, ChangeSetName=name, TemplateURL=template_url,
Parameters=[{"ParameterKey": k, "ParameterValue": v} for k, v in params.items()],
Capabilities=["CAPABILITY_NAMED_IAM"], ChangeSetType="UPDATE")
try:
cfn.get_waiter("change_set_create_complete").wait(StackName=stack, ChangeSetName=name)
except Exception:
reason = cfn.describe_change_set(StackName=stack, ChangeSetName=name)["StatusReason"]
if "didn't contain changes" in reason:
return "no-op"
raise
risky = []
for page in cfn.get_paginator("describe_change_set").paginate(StackName=stack, ChangeSetName=name):
for ch in page["Changes"]:
rc = ch["ResourceChange"]
if rc["ResourceType"] in STATEFUL and (
rc["Action"] == "Remove" or rc.get("Replacement") in ("True", "Conditional")):
risky.append(f'{rc["Action"]} {rc["LogicalResourceId"]} replacement={rc.get("Replacement")}')
if risky and not approved:
print("refusing to execute:", *risky, sep="\n ")
sys.exit(2)
cfn.execute_change_set(StackName=stack, ChangeSetName=name)
cfn.get_waiter("stack_update_complete").wait(StackName=stack)
Policies that protect data
Several template attributes exist to protect data, and they are worth setting by default on every stateful resource.
- DeletionPolicy controls what happens when the resource is removed from the template or the stack is deleted:
Delete,Retain(leave it in the account, now unmanaged),Snapshotfor resources that support it such as RDS and EBS volumes, andRetainExceptOnCreate, which retains the resource except when the stack operation that created it rolls back, avoiding litter from failed first deployments. - UpdateReplacePolicy applies the same choices to the old physical resource during a replacement. Without it, a replacement deletes the old database after the new empty one is created.
- Stack policies are JSON documents that deny update actions, such as Update:Replace or Update:Delete, on chosen logical IDs during stack updates, a guard against accidental replacement.
- Termination protection blocks deletion of the whole stack until someone explicitly turns it off.
Rollback and stuck stacks
If any resource fails during create or update, CloudFormation by default rolls back: it reverts updated resources to their previous configuration and deletes newly created ones, walking the graph in reverse. A failed first create ends in ROLLBACK_COMPLETE, a state from which the stack can only be deleted and recreated. For development, the --disable-rollback option keeps successful resources so you can fix the failing one and retry instead of waiting for a full rollback and redeploy.
The painful state is UPDATE_ROLLBACK_FAILED: the rollback itself failed, often because someone changed or deleted a resource outside CloudFormation, or a dependency such as a custom resource cannot revert. Fix the underlying cause, then run aws cloudformation continue-update-rollback; if a resource cannot be rolled back, pass its logical ID in --resources-to-skip, which marks it complete so the rollback can finish. The skipped resource's real state may now differ from the template, so reconcile it immediately afterwards.
Signals and rolling updates
A resource reaching CREATE_COMPLETE only means the API call succeeded. An EC2 instance is complete when it is launched, not when its application works. A CreationPolicy makes CloudFormation wait for success signals, sent with the cfn-signal helper or the SignalResource API, within a timeout; if the signals do not arrive, the resource fails and the stack rolls back. For Auto Scaling groups, an UpdatePolicy with AutoScalingRollingUpdate replaces instances in batches and can wait for signals per batch, giving a health-checked rolling deployment.
WebGroup:
Type: AWS::AutoScaling::AutoScalingGroup
CreationPolicy:
ResourceSignal: { Count: 2, Timeout: PT15M }
UpdatePolicy:
AutoScalingRollingUpdate:
MinInstancesInService: 2
MaxBatchSize: 1
WaitOnResourceSignals: true
PauseTime: PT10M
Properties:
MinSize: '2'
MaxSize: '6'
VPCZoneIdentifier: !Ref AppSubnets
LaunchTemplate:
LaunchTemplateId: !Ref WebLaunchTemplate
Version: !GetAtt WebLaunchTemplate.LatestVersionNumber
# In the launch template user data, after the app passes its local health check:
# /opt/aws/bin/cfn-signal -e $? --stack ${AWS::StackName} --resource WebGroup --region ${AWS::Region}
Custom resources, Hooks and Guard
When CloudFormation lacks a resource type, a custom resource fills the gap: CloudFormation sends Create, Update and Delete requests to a Lambda function or SNS topic, and the handler must send a response to a pre-signed S3 URL. A handler that crashes without responding leaves the stack waiting until the timeout, which historically was an hour; set the ServiceTimeout property to fail faster. The response data is limited to 4,096 bytes. Handle Delete correctly, including deletes of resources that failed to create, or stack deletion and rollback get stuck. For reusable types, the CloudFormation registry lets you publish resource types with proper schemas and handlers instead.
For organisation-wide controls, Hooks run checks before resources are provisioned and can fail an operation, and cfn-guard evaluates policy rules against templates in CI. Use both: Guard catches problems in pull requests, Hooks enforce them for everyone, including people deploying from laptops.
Cross-stack wiring, drift and StackSets
Stacks share values in two ways. Cross-stack references export an output by name and consume it with Fn::ImportValue. They are simple but create lock-in: an export cannot be changed or deleted while any stack imports it, so a network stack exporting subnet IDs cannot replace those subnets without first updating every consumer. Prefer SSM Parameter Store for values that may change, read with a parameter of type AWS::SSM::Parameter::Value<String> or a dynamic reference. Nested stacks are child stacks deployed by a parent as resources; they update together and roll back together. A single nested-stack operation can create, update or delete at most 2,500 resources. Use nested stacks for composition within one lifecycle and separate stacks for separate lifecycles and owners.
Dynamic references such as {{resolve:secretsmanager:prod/db:SecretString:password}} pull values at deploy time so secrets never appear in templates. Resources changed outside CloudFormation create drift; drift detection compares live configuration with the template for supported resource types. Run it on a schedule and treat drift as an incident. Existing resources can be brought under management with a resource import change set. StackSets deploy one template to many accounts and Regions, with concurrency and failure tolerance settings, and with service-managed permissions they integrate with AWS Organizations for automatic deployment to new accounts.
Worked example: a rename that deletes data
A team renames a DynamoDB table's logical ID from Orders to OrdersTable during a CDK refactor. The change set shows Add OrdersTable and Remove Orders. Executed blindly, CloudFormation creates a new empty table and deletes the old one with its data. The guarded script above refuses because a stateful resource is being removed. Three fixes exist: revert the logical ID (in CDK, override it), set DeletionPolicy Retain on the old resource, deploy, remove it from the template so it is orphaned intact, then import it under the new logical ID; or use CloudFormation's stack refactoring support, if available to you, which moves resources between logical IDs without replacement. The team adds DeletionPolicy and UpdateReplacePolicy Retain to every stateful resource so the next mistake fails safe.
Failure modes
- Logical ID renames create a new resource and delete the old one, data included, unless a retain policy stops it.
- Custom-named resources needing replacement fail the update because the new name collides with the old.
- Out-of-band changes cause drift, then surprise diffs, failed updates or UPDATE_ROLLBACK_FAILED.
- Custom resources that never respond hang the stack until the timeout, and a broken Delete handler blocks stack deletion.
- Locked exports stop a producer stack from changing a value while any stack imports it.
- Missing signals mark broken instances complete, or a too-short signal timeout rolls back healthy deployments.
- Failed first creates leave ROLLBACK_COMPLETE stacks that must be deleted before redeploying.
Trade-offs
CloudFormation is AWS-native, free to use for AWS resource types, keeps state on the service side so there is no state file to lock or lose, and supports new services through the registry, though sometimes later than the API itself. Against Terraform it trades multi-cloud reach and a richer planning experience for managed state, automatic rollback and deep integration with StackSets, Service Catalog and Control Tower. Rollback is both a feature and a cost: safe by default, slow when deployments fail repeatedly. Raw YAML scales poorly for large systems; CDK adds abstraction but hides logical IDs, so review the synthesised template's change set rather than only the code diff.
What to do next
- Set DeletionPolicy and UpdateReplacePolicy on every stateful resource.
- Enable termination protection on production stacks.
- Deploy only through change sets, and block stateful replacement or removal automatically.
- Let CloudFormation generate names for resources that might ever need replacement.
- Add CreationPolicy and signals for anything with an application inside.
- Run cfn-lint and cfn-guard in CI; enforce key rules with Hooks.
- Replace brittle exports with SSM parameters for values that may change.
- Schedule drift detection and practise recovering from UPDATE_ROLLBACK_FAILED.
- Go further with the AWS CDK, IAM, AWS Config, Control Tower and Terraform.