Google Cloud Deploy is a managed continuous delivery service. It takes a container image that your CI system already built, renders Kubernetes or Cloud Run manifests for each environment, and moves that exact rendered output through an ordered sequence of targets: dev, then staging, then production, with approvals, verification jobs, canary phases and rollback recorded along the way. It does not build or test code. Its job starts at the point where you have an immutable artifact and need to decide, environment by environment, whether it is allowed to run.
This article explains the resource model from first principles, walks through a complete pipeline for a Cloud Run service with a canary in production, shows the automation and policy resources that remove manual clicks without removing control, and lists the failure modes teams actually hit. Building the image is covered in Cloud Build; the runtimes themselves are covered in Cloud Run and GKE.
The resource model
Cloud Deploy has a small vocabulary, and most confusion comes from mixing up two of its words: a release and a rollout. A release is an immutable snapshot of what you intend to deploy: the image references, the Skaffold configuration and the manifests rendered from it for every target in the pipeline. A rollout is one attempt to apply one release to one target. Promoting a release to the next stage creates a new rollout; it does not re-render anything. That is the property the service is built around: what passed staging is byte-for-byte what reaches production, apart from the per-target differences you declared up front with Skaffold profiles.
| Resource | What it is | Lifetime |
|---|---|---|
| DeliveryPipeline | Ordered list of stages; each stage names a target, optional profiles and a strategy (standard or canary) | Long-lived, declared in YAML |
| Target | Where to deploy: a GKE cluster, a Cloud Run location, an attached cluster, or a custom target type; plus requireApproval and execution settings | Long-lived |
| Release | Immutable bundle: images, Skaffold source, rendered manifests per target | One per build you intend to ship |
| Rollout | Application of a release to one target; contains phases, which contain jobs (deploy, verify, predeploy, postdeploy) | One per promotion |
| Automation | Rules that promote releases, advance canary phases or repair failed rollouts without a human | Long-lived, per pipeline |
| DeployPolicy | Restrictions such as freeze windows that block rollouts on selected targets | Long-lived |
Rendering and deploying are done by Skaffold running inside Cloud Deploy's execution environment, which by default runs on Cloud Build workers using a service account you configure per target. You never operate those workers, but you do own their permissions, and that is where many first deployments fail.
Architecture and data flow
Follow the data. CI pushes an image and records its digest. CI then calls gcloud deploy releases create with that digest. Cloud Deploy stores the Skaffold source as a tarball in a Cloud Storage bucket, runs a render job for each target in the pipeline, and stores the rendered manifests alongside the release. The first rollout to the first target starts automatically. From then on, every promotion reads the stored manifests rather than your repository, so a commit that lands after the release was created cannot leak into production through this pipeline.
Worked example: a three-stage Cloud Run pipeline
Here is a complete configuration for a service called web on Cloud Run, with three targets and a canary in production. Pipelines and targets are declared in one file, conventionally clouddeploy.yaml, and registered with gcloud deploy apply.
apiVersion: deploy.cloud.google.com/v1
kind: DeliveryPipeline
metadata:
name: web
serialPipeline:
stages:
- targetId: dev
profiles: [dev]
- targetId: staging
profiles: [staging]
strategy:
standard:
verify: true
- targetId: prod
profiles: [prod]
strategy:
canary:
runtimeConfig:
cloudRun:
automaticTrafficControl: true
canaryDeployment:
percentages: [10, 50]
verify: true
---
apiVersion: deploy.cloud.google.com/v1
kind: Target
metadata:
name: dev
run:
location: projects/acme-web/locations/us-central1
---
apiVersion: deploy.cloud.google.com/v1
kind: Target
metadata:
name: staging
run:
location: projects/acme-web/locations/us-central1
---
apiVersion: deploy.cloud.google.com/v1
kind: Target
metadata:
name: prod
requireApproval: true
run:
location: projects/acme-prod/locations/us-central1The Skaffold file tells Cloud Deploy how to render manifests and what the verification container is. Profiles swap in per-environment Cloud Run service YAML, for example different minimum instances or environment variables, while the image reference stays a placeholder that Cloud Deploy fills from the release.
apiVersion: skaffold/v4beta7
kind: Config
manifests:
rawYaml: [run/service-dev.yaml]
deploy:
cloudrun: {}
verify:
- name: smoke
container:
name: smoke
image: us-docker.pkg.dev/acme-ci/tools/smoke:1.3
command: ["/bin/smoke"]
args: ["--path=/healthz", "--path=/api/v1/ping"]
profiles:
- name: dev
manifests:
rawYaml: [run/service-dev.yaml]
- name: staging
manifests:
rawYaml: [run/service-staging.yaml]
- name: prod
manifests:
rawYaml: [run/service-prod.yaml]With a canary of [10, 50], the production rollout has three phases: canary-10, canary-50 and stable. Each canary phase deploys the new revision, shifts that share of traffic to it, runs the verify container, and then waits to be advanced. Pin the Skaffold schema version to one your Cloud Deploy release supports; the service documents supported Skaffold versions and you can set one explicitly on the release.
Releasing, promoting and rolling back
The day-to-day commands map directly onto the resource model. Use image digests, never mutable tags, so the release really is immutable.
# Register or update the pipeline and targets
gcloud deploy apply --file=clouddeploy.yaml --region=us-central1
# CI: create a release from the digest it just pushed
gcloud deploy releases create web-1-4-2 \
--delivery-pipeline=web --region=us-central1 \
--skaffold-file=skaffold.yaml \
--images=web=us-docker.pkg.dev/acme-ci/app/web@sha256:3f1c...
# Promote to the next target after dev looks healthy
gcloud deploy releases promote --release=web-1-4-2 \
--delivery-pipeline=web --region=us-central1
# A human with the approver role approves the prod rollout
gcloud deploy rollouts approve web-1-4-2-to-prod-0001 \
--release=web-1-4-2 --delivery-pipeline=web --region=us-central1
# Move from canary-10 to the next phase after reading dashboards
gcloud deploy rollouts advance web-1-4-2-to-prod-0001 \
--release=web-1-4-2 --delivery-pipeline=web --region=us-central1
# Something is wrong: redeploy the previous successful release
gcloud deploy targets rollback prod --delivery-pipeline=web --region=us-central1Rollback deserves a precise mental model. It creates a new rollout of an earlier release, using that release's stored manifests. It does not undo anything your new code did outside the manifests: a database migration that dropped a column stays dropped. Expand-and-contract schema changes, where each release can run against both the old and the new schema, are what make rollback safe.
Canaries and verification
Canary mechanics differ by runtime, and the difference matters when you choose percentages.
- Cloud Run. Revisions and traffic splitting are native, so a 10 percent canary is a real 10 percent of requests to the new revision, independent of instance counts. With
automaticTrafficControl: true, Cloud Deploy manages the traffic split for you. - GKE with service networking. Cloud Deploy creates a second Deployment for the canary behind the same Service and sizes it relative to the stable Deployment. Traffic share is therefore approximated by pod counts: with four replicas you cannot get a true 10 percent, and the effective share depends on how the Service balances connections. Long-lived connections such as gRPC streams make it worse.
- GKE with Gateway API. Configuring
gatewayServiceMeshwith an HTTPRoute lets Cloud Deploy set route weights, which gives request-level percentages that do not depend on replica counts.
Verification is a container you supply. It runs as a job inside the phase and its exit code is the verdict: zero passes, anything else fails the phase. Keep it short and deterministic, aim it at real user paths, and make it query your metrics backend for the canary's error rate rather than only hitting a health endpoint that returns 200 while the business logic is broken. Newer pipeline configurations can also declare verify containers directly as verify: tasks in the stage instead of the Skaffold stanza. For work before or after a deploy, such as warming a cache or posting a change record, Cloud Deploy supports predeploy and postdeploy hooks that run Skaffold custom actions defined in skaffold.yaml.
Automation and deploy policies
Manual promotion is the right starting point, but a pipeline where someone clicks promote fifty times a week trains people to click without looking. Automation resources move the routine steps to rules and leave humans the decisions that matter. An automation needs its own service account; it does not fall back to a default one.
apiVersion: deploy.cloud.google.com/v1
kind: Automation
metadata:
name: web/routine
serviceAccount: deploy-automation@acme-ci.iam.gserviceaccount.com
selector:
targets:
- id: "*"
rules:
- promoteReleaseRule:
id: promote-to-next
wait: 15m
toTargetId: "@next"
- advanceRolloutRule:
id: advance-canary
sourcePhases: ["canary-10"]
wait: 20m
- repairRolloutRule:
id: retry-then-rollback
repairPhases:
- retry:
attempts: 2
wait: 1m
- rollback: {}Read that as policy. The selector applies to the whole automation, not to each rule, so with id: "*" a release that succeeds on any target is promoted to the next one after fifteen minutes of soak: dev to staging, and staging to prod. Production is held only because its target requires approval; automation does not bypass requireApproval. If you want promotion to stop at staging, split this into two automations with narrower selectors. Once approved, the canary-10 phase advances itself after twenty quiet minutes, and a failed job is retried twice, then rolled back. Restrict the selector if you do not want repair rules rolling back production on their own, and check field names against the configuration schema for your region's release of the service, since automation rules have gained options over time.
Deploy policies handle the opposite need: stopping rollouts. A policy with a weekly time window can block production rollouts on Friday evenings or across a holiday freeze. Overriding a policy is possible for an emergency fix, but it requires an explicit flag on the command and a dedicated IAM permission, so the override is visible and auditable rather than routine.
apiVersion: deploy.cloud.google.com/v1
kind: DeployPolicy
metadata:
name: holiday-freeze
selectors:
- target:
id: prod
rules:
- rolloutRestriction:
id: no-weekend-deploys
timeWindows:
timeZone: America/New_York
weeklyWindows:
- daysOfWeek: [SATURDAY, SUNDAY]
startTime: "00:00"
endTime: "24:00"
Identities and execution environments
Three identities are involved and they should be different. The caller, usually your CI service account, needs a releaser role on the pipeline to create releases, plus iam.serviceAccountUser on the execution service account it causes to run. Approvers are humans or a group with the approver role, which lets them approve rollouts without being able to edit pipelines. The execution service account configured on each target is what actually renders and deploys; it needs deploy rights on the runtime (for example Cloud Run developer on the production project) and read access to the artifact registry. See GCP IAM for how these grants compose.
If you leave the execution environment unset, Cloud Deploy uses the Compute Engine default service account, which in many projects is far too powerful. Create one execution account per environment, grant production rights only to the production one, and put production targets in a separate project so a compromised dev pipeline cannot reach them. If your clusters are private, configure a private Cloud Build worker pool in the target's execution settings so deploy jobs can reach the control plane.
Failure modes
| Symptom | Usual cause | Fix |
|---|---|---|
| Release creation fails at render | Skaffold schema version unsupported, or profile names do not match the pipeline | Pin a supported Skaffold version; render locally with the same profile |
| Deploy job fails with permission denied | Execution service account lacks runtime or registry rights in the target project | Grant per-target roles; read the job's Cloud Build log, not just the rollout status |
| Production runs a different image than staging | Release used a mutable tag that was re-pushed | Pass digests in --images; reject tags in CI |
| Canary looks healthy, full rollout fails | Canary share too small to surface errors, or verify only hits a health endpoint | Verify against error-rate metrics; size the canary to get enough requests |
| Rollback succeeds but users still see errors | Data or schema change is not backward compatible | Expand-and-contract migrations; separate migration releases |
| Cluster drifts from the pipeline | Someone ran kubectl apply by hand | Restrict direct deploy rights; treat Cloud Deploy as the only writer |
| Approvals pile up and get rubber-stamped | Every target requires approval | Approve only production; automate lower environments |
Trade-offs
Cloud Deploy is a push model: a central service applies changes to targets. Pull-based GitOps tools such as Argo CD or Flux run inside each cluster and reconcile it continuously against a Git repository. Pull models correct drift automatically and keep cluster credentials inside the cluster; Cloud Deploy gives you a managed audit trail, approvals, canaries for Cloud Run as well as GKE, and nothing to operate. Teams that are mostly on Cloud Run usually find Cloud Deploy the shorter path, because there is no cluster to host a GitOps controller.
The simplest alternative is a deploy step at the end of a Cloud Build pipeline. That is fine for one environment, but it re-runs the build context for each environment, has no promotion history and no approval object. The point at which a team starts asking which commit is in production, and who allowed it, is usually the point to move to a delivery pipeline. For infrastructure rather than application rollout, use a separate tool; Deployment Manager explains the infrastructure-as-code side and its migration path.
What to do next
- Write
clouddeploy.yamlwith dev, staging and prod targets, and put prod in its own project. - Create one execution service account per target; remove the Compute default account from the path.
- Change CI to create releases with image digests and to fail if a tag is passed.
- Add a verify container that checks real endpoints and the canary's error rate.
- Turn on a canary for production and choose percentages that give enough traffic to judge.
- Add an automation that promotes dev to staging and retries flaky jobs; keep prod approval manual at first.
- Add a deploy policy for your freeze windows and rehearse an override once.
- Run a rollback drill in staging, including a release that carries a schema change.