Every build system eventually needs a place to put what it builds: container images, Python wheels, npm tarballs, Maven jars, OS packages, Helm charts. On Google Cloud that place is Artifact Registry, the managed, multi-format successor to Container Registry. It looks like a simple bucket with a package protocol in front, and for a weekend project it is. In a real organization it becomes a supply-chain control point: what you allow in from the public internet, who can push, whether a tag can be moved, what gets scanned, what may be deployed and how long anything is kept.
This article explains the resource model from first principles, the three repository modes, how each client authenticates, a worked pipeline from source to a digest-pinned deployment, the access and network controls worth turning on, cleanup policies that actually save money, and the failure modes teams hit in production. Commands use the current gcloud CLI; flags change, so confirm against the documentation for your gcloud version before scripting them.
What it is and what it replaced
Container Registry stored images in Cloud Storage buckets behind gcr.io hostnames, with permissions managed on the buckets. It was deprecated, stopped accepting writes on 18 March 2025, and stopped serving reads from its own storage in June 2025. Artifact Registry replaced it with first-class repositories that have their own IAM policies, formats beyond Docker and regional placement you choose. Projects that set up gcr.io repositories in Artifact Registry keep their gcr.io URLs, now served by Artifact Registry, so existing manifests did not need rewriting.
Supported formats include Docker and OCI images, which also covers Helm charts stored as OCI artifacts, Maven, npm, Python, Go modules, Apt and Yum packages, Kubeflow pipeline templates and a generic format for arbitrary files. A repository holds exactly one format, which is the first design decision you make.
The resource model: versions versus tags
The hierarchy is project, location, repository, package, version, tag. A location is a region such as us-central1 or a multi-region such as us or europe. The location is part of the hostname, which is why moving a repository between regions is a migration, not a setting. A Docker image reference looks like us-central1-docker.pkg.dev/acme-prod/app-images/api:1.4.2: region and format in the host, then project, repository and image.
Inside a repository a package is a named thing such as an image or a Python distribution; a version is one immutable piece of content, identified for images by its sha256 digest; a tag is a mutable, human-friendly pointer to a version. The distinction between version and tag is the most important idea in the service. A tag like latest or 1.4.2 can be moved to point at different bytes tomorrow; a digest cannot. Every safe practice below comes back to deploying digests and treating tags as labels.
Standard, remote and virtual repositories
Repositories come in three modes. A standard repository stores artifacts you push. A remote repository is a pull-through cache of an upstream such as Docker Hub, Maven Central, npmjs or PyPI: the first request fetches from upstream and stores a copy, later requests are served from the cache. That protects builds from upstream outages and rate limits, keeps pulls inside Google's network, and gives you a single point to scan and audit everything that enters from the internet. For Docker Hub, configure the remote with credentials stored in Secret Manager so cache fills count against an authenticated limit rather than an anonymous one.
A virtual repository holds nothing itself. It presents one URL that fans out to an ordered list of upstream repositories in the same location and format, each with a priority. Clients configure one index; the virtual repository resolves a package from the highest-priority upstream that has it. This is also your defence against dependency confusion, where an attacker publishes a public package with the same name as an internal one: give the internal standard repository the higher priority so an internal name never resolves from the public cache.
# Standard repository for release images, tags cannot be moved once set
gcloud artifacts repositories create app-images \
--repository-format=docker --location=us-central1 \
--immutable-tags --description="Release images"
# Pull-through cache of Docker Hub and of PyPI
gcloud artifacts repositories create dockerhub \
--repository-format=docker --location=us-central1 \
--mode=remote-repository --remote-docker-repo=DOCKER-HUB
gcloud artifacts repositories create pypi-cache \
--repository-format=python --location=us-central1 \
--mode=remote-repository --remote-python-repo=PYPI
# One Python index: internal packages win over the public cache
cat > upstreams.json <<'EOF'
[{"id": "internal", "repository": "projects/acme-prod/locations/us-central1/repositories/py-internal", "priority": 100},
{"id": "pypi", "repository": "projects/acme-prod/locations/us-central1/repositories/pypi-cache", "priority": 10}]
EOF
gcloud artifacts repositories create py \
--repository-format=python --location=us-central1 \
--mode=virtual-repository --upstream-policy-file=upstreams.json
Authenticating every client
Artifact Registry speaks each format's native protocol, so the work is getting short-lived Google credentials into each client. Never paste a service-account key into a config file; every client has a helper that mints tokens from Application Default Credentials, which on Google Cloud compute come from the attached service account and on a laptop from gcloud auth application-default login.
# Docker: register gcloud as a credential helper for this regional host
gcloud auth configure-docker us-central1-docker.pkg.dev
# Python: keyring backend fetches tokens automatically
pip install keyring keyrings.google-artifactregistry-auth
pip install --index-url https://us-central1-python.pkg.dev/acme-prod/py/simple/ acme-utils
# npm: write the registry in .npmrc, then refresh the token before installs
# @acme:registry=https://us-central1-npm.pkg.dev/acme-prod/npm-internal/
# //us-central1-npm.pkg.dev/acme-prod/npm-internal/:always-auth=true
npx google-artifactregistry-auth
# Print ready-made client settings for any repository
gcloud artifacts print-settings python --repository=py --location=us-central1GKE nodes and Cloud Run pull with their service account and need no client configuration, only the reader role on the repository. A frequent surprise is the node service account on a GKE cluster in another project: grant it roles/artifactregistry.reader on the repository explicitly, or pods sit in ImagePullBackOff with a 403 in the events.
Worked example: commit to digest-pinned deploy
Here is a complete path from commit to production for a service called api. The build pushes an image tagged with the commit SHA, resolves the digest that tag points at, and deploys the digest. Because the repository has immutable tags, a second push of the same SHA tag fails loudly instead of silently replacing what is running.
# cloudbuild.yaml
substitutions:
_IMAGE: us-central1-docker.pkg.dev/acme-prod/app-images/api
steps:
- name: gcr.io/cloud-builders/docker
args: ["build", "-t", "${_IMAGE}:${SHORT_SHA}", "."]
- name: gcr.io/cloud-builders/docker
args: ["push", "${_IMAGE}:${SHORT_SHA}"]
- name: gcr.io/google.com/cloudsdktool/cloud-sdk
entrypoint: bash
args:
- -c
- |
DIGEST=$$(gcloud artifacts docker images describe "${_IMAGE}:${SHORT_SHA}" \
--format='value(image_summary.digest)')
echo "deploying ${_IMAGE}@$${DIGEST}"
gcloud run deploy api --region=us-central1 --image="${_IMAGE}@$${DIGEST}"The doubled dollar signs matter: Cloud Build substitutes $NAME and ${NAME} itself, so shell variables and command substitution must be escaped as $$. The build's service account needs roles/artifactregistry.writer on app-images and permission to deploy; the Cloud Run service agent needs reader. The same digest can then be promoted to staging and production by reference, without rebuilding, which is what makes the artifact you tested the artifact you ship. For GKE, write the digest into the manifest with Kustomize or Helm in the same step, and record it in the release notes.
Access and network controls
Repositories have their own IAM policies, so grant at the repository, not the project. The useful roles are reader for runtimes and developers, writer for exactly one build identity per repository, repoAdmin for the platform team, which adds deleting versions, and admin for creating repositories. A writer role held by a human is a red flag: anyone who can push to a release repository can replace production code.
gcloud artifacts repositories add-iam-policy-binding app-images \
--location=us-central1 \
--member=serviceAccount:builder@acme-prod.iam.gserviceaccount.com \
--role=roles/artifactregistry.writerBeyond IAM, four controls matter. Immutable tags stop a moved tag from rewriting history. VPC Service Controls put the Artifact Registry API inside a perimeter, so a stolen credential cannot pull images from outside your networks; private nodes then reach the registry through Private Google Access. Customer-managed encryption keys are available when a policy demands them, chosen at repository creation and not changeable afterwards. Data Access audit logs, off by default, record who pulled what, which you want for release repositories and will need for cleanup dry runs.
Cleanup policies that do not break production
Without cleanup a busy repository grows forever, and storage is billed per gigabyte-month. A service that builds fifty 400 MB images a day accumulates about 7 TB a year, mostly images nobody will run again; layers shared between images are stored once, so the real figure is lower but the curve is the same. Cleanup policies are JSON rules attached to a repository. Each has a delete or keep action and conditions on tag state, tag prefixes, version name prefixes, package name prefixes and age. A keep action can also keep the most recent N versions per package. When a version matches both a delete and a keep policy, it is kept.
[
{"name": "drop-untagged-after-7d", "action": {"type": "Delete"},
"condition": {"tagState": "untagged", "olderThan": "7d"}},
{"name": "drop-ci-tags-after-30d", "action": {"type": "Delete"},
"condition": {"tagState": "tagged", "tagPrefixes": ["ci-", "pr-"], "olderThan": "30d"}},
{"name": "keep-releases", "action": {"type": "Keep"},
"condition": {"tagState": "tagged", "tagPrefixes": ["v"]}},
{"name": "keep-last-20", "action": {"type": "Keep"},
"mostRecentVersions": {"keepCount": 20}}
]gcloud artifacts repositories set-cleanup-policies app-images \
--location=us-central1 --policy=policy.json --dry-run
# after reviewing the audit logs for a day or two
gcloud artifacts repositories set-cleanup-policies app-images \
--location=us-central1 --policy=policy.json --no-dry-runPolicies run in a periodic background job, and changes take effect within about a day, so do not expect instant results. Always start in dry-run mode, then read the Data Access audit logs to see what would have been deleted. The trap is the image a production workload runs by digest with no tag: it is untagged, so the first rule deletes it, and the next node that scales up cannot pull. Tag every deployed digest with a protected prefix, or derive a keep list from what is actually running before you switch dry run off. Images referenced by a multi-architecture index are not deleted while the parent manifest exists.
Scanning and deploy-time enforcement
Turning on the Container Scanning API makes Artifact Analysis scan images on push and continuously re-evaluate recently pushed or pulled images as new vulnerabilities are published, covering OS packages and common language package managers. The On-Demand Scanning API scans an image before it is pushed, which is useful as a CI gate. Scan results are only advice; enforcement belongs at deploy time. Binary Authorization can require that an image carry an attestation, for example that it was built by your Cloud Build pipeline and passed a vulnerability threshold, before GKE or Cloud Run will run it. That closes the gap where someone with deploy rights runs an image straight from a remote cache.
Placement, latency and cost
Place repositories in the region where most pulls happen. Pulls from a repository by compute in the same region avoid inter-region network charges and are fastest; a multi-region location trades a little latency for broader availability. For a global fleet, replicate release images to regional repositories from the pipeline rather than pulling across continents on every node start. Large machine-learning images, often several gigabytes with CUDA libraries, make this visible: a GKE node pool scaling from zero pulls the image on every new node, so registry placement shows up directly in scale-up latency. Image streaming on GKE reduces the wait for supported images, and keeping base layers stable means most pulls fetch only the small top layers. Check current storage and network pricing on the pricing page rather than hard-coding numbers into cost models.
Failure modes
The failures teams actually hit:
- Deploying tags. A manifest says
:latest; a push changes what new nodes run while old nodes keep the previous image. Two versions serve traffic and nobody knows. Deploy digests. - 403 on pull from another project. The runtime service account lacks reader on the repository. Check which identity the node or service actually runs as.
- Cleanup deletes a running image. Untagged digest pins match an untagged-delete rule. Dry run first and protect deployed digests with tags.
- Dependency confusion through a virtual repository. Public cache has higher priority than the internal repository, or internal names are not reserved. Order priorities deliberately and test with a canary package name.
- Docker Hub limits during cache fills. An unauthenticated remote repository hits anonymous limits on a cold cache. Configure credentials.
- Stale npm tokens. Tokens written by the helper expire in about an hour; CI jobs that cache .npmrc fail mid-week. Refresh tokens at job start.
- Region baked into every reference. Repositories cannot move; a region change touches every manifest. Centralize the registry host in one variable.
Trade-offs
| Decision | Option A | Option B | Guidance |
|---|---|---|---|
| Repo granularity | One per team | One per trust boundary | Split by who may push, not by org chart |
| Location | Region | Multi-region | Region near compute; replicate for global fleets |
| Public deps | Pull from internet | Remote repo cache | Cache: availability, audit, scanning |
| Tags | Mutable | Immutable | Immutable for release repos; mutable only for scratch |
| Retention | Keep all | Cleanup policies | Policies with dry run and protected prefixes |
| Enforcement | Scan only | Scan plus Binary Authorization | Enforce at deploy for production |
What to do next
Work through this list in order:
- Inventory existing repositories, their locations and who holds writer; remove human writers.
- Create remote repositories for every public registry you depend on and a virtual repository with internal upstreams at the highest priority.
- Turn on immutable tags for release repositories and change every deploy to reference a digest.
- Enable Data Access audit logs, write a cleanup policy, and run it in dry-run mode for a week.
- Enable vulnerability scanning and add a Binary Authorization policy for production clusters.
- Put the API inside a VPC Service Controls perimeter if you handle regulated data.
- Continue with Cloud Build for the pipeline side, GKE and Cloud Run in depth for the runtimes that pull these images, GCP IAM for least-privilege design, and Security Command Center to see vulnerability findings centrally.