Every production system has credentials: database passwords, API keys for payment providers, signing keys, TLS private keys. They end up in environment files, CI variables, Kubernetes Secrets and, eventually, in a Git history. Secret Manager is Google Cloud's managed store for them: a versioned, access-controlled, audited home for small sensitive blobs, reached through an API that checks IAM on every read.

Using the API is easy. Running it well is where teams struggle: choosing replication, deciding who may read what, getting new values to running workloads, rotating without an outage, and not melting the quota when a thousand pods start at once. This article covers the model, the access path, consumption on each compute platform, rotation, lifecycle, limits and failure modes, with code and a checklist.

Advertisement

The model: secrets, versions, aliases

A secret is a named container, projects/P/secrets/orders-db-password, carrying metadata: replication policy, labels, annotations, optional expiry, rotation schedule and notification topics. It holds no value itself. Each value is a version, numbered 1, 2, 3 in order of creation. Versions are immutable: to change a password you add a version, never edit one. A payload can be up to 64 KiB, which fits passwords, keys and certificates but not configuration files; for those, Google's separate Parameter Manager allows 1 MiB per version.

Each version is ENABLED, DISABLED or DESTROYED. Disabled versions cannot be accessed but can be re-enabled; destroyed versions lose their payload permanently. The alias latest always resolves to the most recently created version, and you can define up to 50 custom aliases per secret, such as current and previous, to point at specific versions.

The latest alias is the most important design decision you make. Reading latest means a new version takes effect as soon as clients re-read, with no deployment. Reading a pinned number means a change requires a deploy, so a bad value can be rolled back by redeploying the previous revision. Most teams pin for anything read once at start-up and use latest with a refresh loop for rotated credentials.

Where the bytes live: replication and regional secrets

A global secret has a replication policy chosen at creation. Automatic lets Google choose where replicas live and is the simplest option. User-managed lists the regions explicitly, which is what data-residency requirements usually demand. Both are reached through the global endpoint. Treat the replication policy as fixed once chosen; plan it rather than expecting to migrate later.

Regional secrets are a separate resource type under projects/P/locations/L/secrets/S, stored in and served from one region through a regional endpoint. Choose them when a secret must never leave a region, or when you need their much higher write limits, covered below. Payloads are encrypted at rest with Google-managed keys by default; with customer-managed encryption keys (CMEK) the payload is encrypted with a Cloud KMS key you control, so disabling that key makes the versions unreadable, which is both the point and a new way to cause an outage.

Advertisement

The access path

One AccessSecretVersion call, and the rotation loop beside itWorkloadCloud Run, GKE, VMMetadata serverservice account tokenADCSecret Manager APIglobal or regional endpointaccess latest / v3IAM checksecretAccessor on secretVersion storereplicas: automatic or listed regionsEncryptionGoogle-managed key or CMEK (Cloud KMS)Cloud Audit LogsAdmin Activity; Data Access if onRotation schedulerotation_periodPub/Sub topicSECRET_ROTATE, VERSION_ADDYour rotatornew credential, add versionTarget systemDB, vendor APISecret Manager only announces that rotation is due; your subscriber does the rotating
Top: the workload's service account token, an IAM check on the secret, then a read and decrypt of the version. Bottom: rotation, where Secret Manager publishes a notification and your code does the work.

A workload never holds a Secret Manager credential. On Cloud Run, GKE with Workload Identity Federation and Compute Engine, Application Default Credentials fetch a short-lived token for the workload's service account from the metadata server. The API checks that identity against IAM on the secret, reads the requested version, decrypts it and returns it. The payload travels over TLS and the response carries a CRC32C checksum for end-to-end verification.

from google.cloud import secretmanager
import google_crc32c

client = secretmanager.SecretManagerServiceClient()
project = "shop-prod"
secret = client.secret_path(project, "orders-db-password")

def crc32c(data: bytes) -> int:
    c = google_crc32c.Checksum()
    c.update(data)
    return int(c.hexdigest(), 16)

def add_version(value: str) -> str:
    data = value.encode("utf-8")
    v = client.add_secret_version(request={
        "parent": secret,
        "payload": {"data": data, "data_crc32c": crc32c(data)},   # server verifies it
    })
    return v.name                                     # .../versions/7

def access(version: str = "latest") -> str:
    resp = client.access_secret_version(request={"name": f"{secret}/versions/{version}"})
    if resp.payload.data_crc32c != crc32c(resp.payload.data):
        raise RuntimeError("payload corrupted in transit")
    return resp.payload.data.decode("utf-8")

# Regional secret: same client class, regional endpoint, locations/ in the path.
regional = secretmanager.SecretManagerServiceClient(
    client_options={"api_endpoint": "secretmanager.europe-west1.rep.googleapis.com"})
name = f"projects/{project}/locations/europe-west1/secrets/orders-db-password/versions/latest"

Access control that holds up

Grant at the level of the individual secret, not the project. A project-level roles/secretmanager.secretAccessor binding lets that identity read every secret in the project, today's and every future one. The general mechanics of bindings and inheritance are in GCP IAM.

RoleCan doGive it to
roles/secretmanager.secretAccessorread payloadsthe one workload that needs that secret
roles/secretmanager.viewerlist and read metadata, not payloadsdashboards, inventory tooling
roles/secretmanager.secretVersionAdderadd versions onlya CI job that writes but should not read (a rotator that reads latest also needs secretAccessor)
roles/secretmanager.secretVersionManageradd, enable, disable, destroy versionsoperators running rotation
roles/secretmanager.admineverything, including IAM on secretsa small platform group, ideally via just-in-time elevation

One service account per workload keeps these grants meaningful; a shared default compute service account means every workload can read every secret any of them can. Use IAM conditions on resource name prefixes when you manage many secrets per team. For exfiltration risk from a stolen token, put Secret Manager inside a VPC Service Controls perimeter so the API refuses calls from outside your networks and projects.

Admin Activity audit logs, which record creating secrets, adding versions and changing IAM, are always on. Reads of payloads are Data Access logs, which most Google Cloud services leave off by default; enable them for Secret Manager, or you cannot answer who read this password last Tuesday.

Getting values into workloads

# Create, add a version from stdin (never from argv or shell history), grant one identity.
gcloud secrets create orders-db-password --replication-policy=automatic
printf '%s' "$NEW_PASSWORD" | gcloud secrets versions add orders-db-password --data-file=-
gcloud secrets add-iam-policy-binding orders-db-password \
  --member=serviceAccount:orders-api@shop-prod.iam.gserviceaccount.com \
  --role=roles/secretmanager.secretAccessor

# Cloud Run: env var pinned to a version; file mount follows latest.
gcloud run deploy orders-api --image=IMAGE \
  --update-secrets=DB_PASSWORD=orders-db-password:7,/secrets/vendor-key=vendor-api-key:latest

Cloud Run resolves secrets exposed as environment variables once, at instance start-up, so Google recommends pinning them to a version. Secrets mounted as files are fetched from Secret Manager when read, so a mounted latest follows rotation without a redeploy. The service identity needs secretAccessor on each secret. Platform details are in Cloud Run architecture.

GKE offers the Secret Manager add-on, a CSI driver (secrets-store-gke.csi.k8s.io) configured with a SecretProviderClass that mounts secrets as files, authenticating as the pod's identity through Workload Identity Federation for GKE. On GKE 1.32.2-gke.1059000 and later it can refresh mounted secrets automatically, with a rotation interval of at least 120 seconds. Syncing into native Kubernetes Secrets, for example with an external operator, is the alternative when an application insists on them, at the cost of copies in etcd. Cluster-side identity is covered in GKE architecture.

VMs and anything else call the API from the application at start-up, cache the value in memory, and refresh on a timer or on failure. Never write the value to disk, logs or crash dumps.

Quotas, caching and the start-up stampede

Access calls are limited to 90,000 per minute per project, about 1,500 per second. Metadata reads and writes are far lower, 600 per minute each. Adding versions to a global secret is limited to 2 per second and 120 per minute per secret; regional secrets allow 80 per second per secret per region.

Work the arithmetic for your fleet. A service with 800 pods that reads 5 secrets per request at 200 requests per second per pod would issue 800,000 access calls per second: over 500 times the quota. The same service reading each secret once at start-up and refreshing every 5 minutes issues 4,000 reads per rollout and about 13 per second in steady state. Cache in process, refresh on a timer with jitter, and treat Secret Manager as a configuration source, never as a per-request dependency. Budget for a full rollout or autoscale event: 800 pods times 5 secrets starting within a minute is 4,000 calls, well within quota, but a crash-looping deployment that restarts every few seconds can consume it.

Rotation: Secret Manager schedules, you rotate

Secret Manager does not change your database password. A rotation schedule, a rotation_period of at least one hour with a next_rotation_time at least five minutes ahead, makes it publish a SECRET_ROTATE message to the Pub/Sub topics configured on the secret, up to 10 of them. Its service agent needs roles/pubsub.publisher on each topic. The same topics receive lifecycle events such as SECRET_VERSION_ADD, SECRET_VERSION_DISABLE and SECRET_VERSION_DESTROY_SCHEDULED; reads never publish anything. Subscription and retry behaviour are covered in Pub/Sub architecture.

# Rotator: a Pub/Sub-triggered function. Two database users (A and B) alternate so the
# credential in use is never the one being changed. Helpers (access_by_name,
# add_version_by_name, db_admin, generate_password) are yours to supply.
import json

def on_secret_event(message):
    attrs = message.attributes
    if attrs.get("eventType") != "SECRET_ROTATE":
        return                                        # ignore VERSION_ADD etc.
    secret_id = attrs["secretId"]                     # projects/.../secrets/orders-db-password
    current = json.loads(access_by_name(f"{secret_id}/versions/latest"))
    standby_user = "orders_b" if current["user"] == "orders_a" else "orders_a"

    new_password = generate_password(32)
    db_admin.alter_user_password(standby_user, new_password)          # 1. change standby
    db_check_login(standby_user, new_password)                        # 2. prove it works
    add_version_by_name(secret_id, json.dumps(
        {"user": standby_user, "password": new_password}))            # 3. publish
    # 4. clients pick it up on their next refresh; the old user stays valid until the
    #    next rotation changes it, which is the grace period.

The two-user pattern matters because clients refresh at different times. If the rotator changed the only password in use, every client would fail from the moment of change until it re-read the secret. With two alternating users, the old credential remains valid for a full rotation period after the new one is published. The rotator must be idempotent, since Pub/Sub may redeliver, and must alert on failure: a rotator that silently fails leaves a credential unrotated indefinitely while the schedule looks healthy.

Version lifecycle and destruction

Old versions accumulate, and each enabled version is a credential someone might still use. After a rotation's grace period, disable the old version rather than destroying it; a client that breaks tells you who was still pinned to it, and re-enabling is instant. Destroy only after a quiet period.

Set a version_destroy_ttl on important secrets to get delayed destruction: a destroy request then disables the version and schedules permanent destruction at scheduled_destroy_time, and an admin can restore it until then. A secret-level expiry, by contrast, deletes the whole secret at a set time, which suits short-lived credentials for temporary environments.

Failure modes

FailureSymptomPrevention
Pinned env var, value rotatedold password keeps working until the DB drops it, then every instance failsfile mounts or refresh loops for rotated secrets
New version is wrongall clients reading latest break togethervalidate before adding; pin critical consumers; disable the bad version to fall back
Access quota exhaustedRESOURCE_EXHAUSTED at start-up during a rolloutcache, jitter, avoid per-request reads
CMEK key disabled or IAM removedevery access failsalert on KMS key state; restrict who can change it
Rotator brokenno error, credential ages past policyalert on time since last SECRET_VERSION_ADD
Value loggedsecret in log storage and exportswrapper types that redact on print; log scanning

What to do next

  1. Inventory where secrets live today: env files, CI variables, Kubernetes Secrets, code. Move them into Secret Manager.
  2. Give each workload its own service account and grant secretAccessor per secret, never per project.
  3. Enable Data Access audit logs for Secret Manager and alert on reads by unexpected principals.
  4. Pin env-var secrets to versions; use file mounts or refresh loops with latest for rotated ones.
  5. Cache in process, refresh with jitter, and check your peak rollout against the access quota.
  6. Add rotation schedules, a Pub/Sub topic and an idempotent rotator using two alternating credentials.
  7. Set version_destroy_ttl on critical secrets and disable old versions before destroying them.
Key takeaway: Secret Manager stores small immutable versions under named secrets, checks IAM on every read and audits what you enable. Pick replication or regional secrets deliberately, grant access per secret per workload, and decide between pinned versions and latest for each consumer. Cache reads instead of calling per request, remember that rotation is a notification your own code must act on, rotate with two alternating credentials, and disable before you destroy.