The article GCP Secret Manager, in depth explains how a single secret works: versions and aliases, replication, per-secret IAM, the access path, quotas, rotation with two alternating credentials, and destroying old versions. Read it first if those are new to you. This article covers the problem that appears once Secret Manager is widely adopted: dozens of projects, hundreds or thousands of secrets, several teams creating them, and nobody able to answer simple questions. Which secrets exist? Who owns each one? Which are still read? Which did Terraform create, and is any plaintext sitting in a state bucket?
The goal is a platform you can operate: a deliberate project layout, infrastructure as code that never copies values into state, organisation guardrails, an inventory of what exists and what is used, and a migration path for secrets still in Kubernetes objects and .env files. Where Google documents a caveat, it is stated.
Where secrets live: projects, names and labels
The first decision is where secrets live, because projects are the unit for organisation policy, VPC Service Controls perimeters, audit configuration and billing. There are three common layouts, and most organisations end up with the second.
| Layout | How it works | Good for | Cost |
|---|---|---|---|
| Co-located | each secret lives in the project of the workload that reads it | small estates, one team per project | inventory and policy are spread across every workload project |
| Secrets project per team and environment | e.g. orders-secrets-prod; workloads in other projects get per-secret access | most organisations | a cross-project IAM grant for every consumer |
| One central secrets project | a platform team owns every secret | very small or highly regulated estates | one project-level mistake exposes everything; the platform team becomes a bottleneck |
A dedicated secrets project per team and environment keeps the blast radius small. A production secret cannot be read through a staging project's permissions, and an owner who leaves can be traced to one project. The project can sit inside a perimeter without moving every workload project into it. Combine this with a naming scheme that makes the purpose of a secret obvious, such as <system>-<purpose> (for example orders-db-password), and labels that carry the facts your reports need. Labels are key-value metadata you can filter on; annotations hold longer free-form notes. Make a small set of labels mandatory and check for them in CI:
owner: a team, never a person, so the label survives staff changes.env: prod, staging or dev, matching the folder.rotation: scheduled, manual or none, which tells the staleness report what to expect.source: which system the credential belongs to (cloudsql, stripe, github-app), for incident response.
Terraform without plaintext in state
Terraform is the natural place to declare secrets and their permissions. The risk is the payload. When a configuration sets a version's value through an ordinary argument, Terraform records it in state, and state is a JSON document in a bucket. Marking a variable sensitive only hides it in CLI output; it does not keep the value out of state. Anyone who can read the state bucket can then read the secret, outside Secret Manager's IAM and audit logs. There are two safe patterns.
Pattern 1: Terraform owns the container, not the value. Terraform creates the secret, its replication, labels, rotation settings and IAM bindings. Versions are added by a rotator, by CI from another secret store, or by a person using a break-glass procedure. This is the right choice for every rotated credential. A Terraform-managed version and a rotator would otherwise fight: each apply would try to restore the version Terraform remembers.
resource "google_secret_manager_secret" "orders_db" {
project = var.secrets_project # orders-secrets-prod
secret_id = "orders-db-password"
labels = {
owner = "team-orders"
env = "prod"
rotation = "scheduled"
source = "cloudsql"
}
replication {
user_managed {
replicas { location = "europe-west1" }
replicas { location = "europe-west4" }
}
}
}
# Per-secret grant to one workload identity: never a project-level binding.
resource "google_secret_manager_secret_iam_member" "orders_api_reads" {
project = var.secrets_project
secret_id = google_secret_manager_secret.orders_db.secret_id
role = "roles/secretmanager.secretAccessor"
member = "serviceAccount:${var.orders_api_sa}"
}Pattern 2: write-only arguments. For values that are set by hand and rarely change, such as a vendor API key, the Google provider's google_secret_manager_secret_version resource accepts secret_data_wo, a write-only argument that is sent to the API but not stored in state, together with secret_data_wo_version. Because Terraform keeps no copy of the value, it cannot tell when the value changes. You increment the version number to make it write a new value. Write-only arguments need recent Terraform and provider releases; confirm both in your pipeline before you rely on them.
variable "vendor_key" {
type = string
sensitive = true # hides it in output; the write-only argument keeps it out of state
}
# google_secret_manager_secret.vendor_key is declared like orders_db above.
resource "google_secret_manager_secret_version" "vendor_key" {
secret = google_secret_manager_secret.vendor_key.id
secret_data_wo = var.vendor_key # supplied as TF_VAR_vendor_key from the CI secret store
secret_data_wo_version = 3 # bump to 4 to push a new value
}With pattern 2 the value still passes through the CI job, so treat its logs and saved plans as sensitive. Either way, search every existing state for secret_data once. Anything found has already leaked to whoever could read the bucket: rotate it.
Organisation guardrails
Organisation policy and IAM defaults stop mistakes before code review has to. Three guardrails cover most of the risk.
Location. The gcp.resourceLocations constraint applies to Secret Manager, with a detail that surprises people. A secret with automatic replication can only be created if the global location is allowed. A secret with user-managed replication can only be created if every location in its replication policy is allowed. It is checked only at creation, so existing secrets are unaffected. Allow only European regions and automatic replication fails, forcing teams to list regions, which is what you want.
Identity. A project-level grant of roles/secretmanager.secretAccessor or broader roles lets an identity read every secret in the project, including future ones. Fail CI on any such binding in a secrets project, and alert when one appears in IAM policy change logs. Grant on the secret, to one service account per workload. The binding mechanics are in GCP IAM.
Perimeter. Put the secrets projects inside a VPC Service Controls perimeter, so a stolen token used from outside your networks is refused even though IAM would allow it. See VPC Service Controls for how perimeters and ingress rules behave. Bring perimeter violations and public-exposure findings into Security Command Center so one team sees them.
Inventory: what exists and what is used
No single source answers both what exists and what is used. Cloud Asset Inventory lists both secretmanager.googleapis.com/Secret and secretmanager.googleapis.com/SecretVersion assets across an organisation, with labels, replication and create times. Google documents two caveats for Secret Manager: the data is synchronised about every seven hours, and change history may be incomplete. Treat it as a daily inventory, not a real-time feed. It also says nothing about reads. Reads appear only in Data Access audit logs, which are off by default and must be enabled for Secret Manager. Before they were enabled, missing reads tell you nothing.
# Daily: export secret metadata for the whole organisation to BigQuery.
gcloud asset export --organization=123456789012 \
--content-type=resource \
--asset-types=secretmanager.googleapis.com/Secret \
--bigquery-table=projects/sec-reporting/datasets/assets/tables/secrets \
--output-bigquery-force
# Once: route Secret Manager Data Access logs from every project to the same dataset.
# Then grant the sink's writer identity BigQuery Data Editor on that dataset.
gcloud logging sinks create sm-reads \
bigquery.googleapis.com/projects/sec-reporting/datasets/audit \
--organization=123456789012 --include-children --use-partitioned-tables \
--log-filter='protoPayload.serviceName="secretmanager.googleapis.com" AND
protoPayload.methodName:"AccessSecretVersion"'The report joins the two sources on the secret's path. Asset names look like //secretmanager.googleapis.com/projects/PROJECT/secrets/ID, and audit resource names end in /versions/N. One side may use the project number and the other the project ID, so inspect a sample row from each table and normalise both before you trust the join. The query below assumes both sides use the same project form.
WITH secrets AS (
SELECT REGEXP_EXTRACT(name, r'projects/[^/]+/secrets/[^/]+') AS path,
JSON_VALUE(resource.data, '$.labels.owner') AS owner,
JSON_VALUE(resource.data, '$.labels.rotation') AS rotation
FROM `sec-reporting.assets.secrets`
),
reads AS (
SELECT REGEXP_EXTRACT(protopayload_auditlog.resourceName,
r'projects/[^/]+/secrets/[^/]+') AS path,
MAX(timestamp) AS last_read,
COUNT(DISTINCT protopayload_auditlog.authenticationInfo.principalEmail) AS readers
FROM `sec-reporting.audit.cloudaudit_googleapis_com_data_access`
WHERE timestamp > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 90 DAY)
GROUP BY path
)
SELECT s.path, s.owner, s.rotation, r.last_read, r.readers
FROM secrets s LEFT JOIN reads r USING (path)
WHERE r.last_read IS NULL OR s.owner IS NULL
ORDER BY s.owner NULLS FIRST, s.path;A secret with no reads in 90 days is a deletion candidate: disable, wait, then destroy. A secret without an owner needs one first. A secret read by many principals is a shared credential; give each consumer its own if the backend allows.
Migrating secrets from Kubernetes and .env files
Most estates already have secrets elsewhere: Kubernetes Secret objects, .env files on VMs, CI variables, sometimes constants in code. Migrate in an order that never leaves a consumer without a value:
- Inventory the old locations, and record for each value its consumers and its owner.
- Create the secret in Terraform (pattern 1), with labels and per-secret IAM for each consumer's identity.
- Copy the value with a script that reads and writes bytes and never prints them.
- Switch each consumer to Secret Manager, pinned to the copied version, and deploy.
- Delete the old copies, including Kubernetes objects, CI variables and files on disk.
- Rotate. The old copies were readable by whoever could read etcd, the CI settings or the VM, so the migrated value is already exposed.
import base64, json, subprocess
from google.cloud import secretmanager
import google_crc32c
client = secretmanager.SecretManagerServiceClient()
def migrate_k8s_secret(namespace: str, name: str, key: str, project: str, secret_id: str) -> str:
"""Copy one key of a Kubernetes Secret into an existing Secret Manager secret. Never prints the value."""
raw = subprocess.run(["kubectl", "get", "secret", name, "-n", namespace, "-o", "json"],
check=True, capture_output=True).stdout
value = base64.b64decode(json.loads(raw)["data"][key]) # bytes, never str-formatted or logged
crc = google_crc32c.Checksum()
crc.update(value)
version = client.add_secret_version(request={
"parent": client.secret_path(project, secret_id), # created by Terraform, pattern 1
"payload": {"data": value, "data_crc32c": int(crc.hexdigest(), 16)},
})
return version.name # pin consumers to this exact version during cut-overThe secret already exists from Terraform, so ownership, labels and IAM come from code, and the returned version name is what the cut-over pins. A set -x in a wrapper script or a debugging print is the usual way migrations leak the values they move.
Worked example: cleaning up 912 secrets
Consider a hypothetical estate: 37 projects, each with secrets created by hand over three years. The first asset export finds 912 secrets. Data Access logging has been on for 120 days, so the 90-day read window is meaningful. The join produces this picture:
| Finding | Count | Action |
|---|---|---|
| no reads in 90 days | 301 | disable latest versions in batches of 50; destroy after 30 quiet days |
| no owner label | 140 | assign by project owner; unowned after 14 days goes on the disable list |
| project-level secretAccessor binding | 9 projects | replace with per-secret grants, then remove |
| read by more than 5 principals | 17 | split into per-consumer credentials where the backend allows |
| value present in Terraform state | 23 | rotate, then move to pattern 1 or 2 |
Disabling is the safety net. When one of the 301 turns out to feed a quarterly job outside the window, the job fails loudly, the version is re-enabled in seconds, and the secret gains an owner and a rotation label. Afterwards about 470 secrets remain in 11 team projects, and the daily report should stay near zero.
Failure modes
| Failure | What you see | Prevention |
|---|---|---|
| Plaintext in Terraform state | nothing, until someone reads the bucket | pattern 1 or 2; scan state; rotate anything found |
| Terraform and a rotator manage the same secret | each apply reverts the rotation, or plans show perpetual drift | Terraform never manages versions of rotated secrets |
| Deletion based on inventory alone | a rarely run job breaks after cleanup | require audit-log evidence; disable before destroying |
| Join on mismatched project forms | every secret looks unused | normalise project number versus ID; check a known-busy secret appears |
| Location policy added after creation | old secrets keep automatic replication | the policy applies at creation; recreate non-compliant secrets deliberately |
| Migration copies left behind | the value is still in etcd or CI settings | delete old copies, then rotate |
Trade-offs
Per-team secrets projects cost more cross-project IAM grants in return for isolation that maps to ownership. Pattern 1 is safe for every secret but needs another way to set initial values; pattern 2 is convenient but needs manual version bumps and routes the value through CI. Data Access logs cost storage and query money, usually little when clients cache, and without them you cannot clean up safely. Strict location policy rules out automatic replication, which for regulated data is the point.
What to do next
- Choose a layout: one secrets project per team per environment, under the matching folder.
- Define the mandatory labels (owner, env, rotation, source) and fail CI when one is missing.
- Move secret declarations into Terraform as containers plus per-secret IAM; use
secret_data_woonly for rarely changed, hand-set values. - Scan every Terraform state for secret payloads and rotate anything you find.
- Set
gcp.resourceLocationsfor regulated folders and block project-level secretAccessor bindings. - Enable Data Access logs for Secret Manager, sink them to BigQuery, and export secret assets daily.
- Run the join weekly: disable unused secrets, assign owners, split widely shared credentials.
- Migrate Kubernetes Secrets and .env files with a non-printing copy script, then delete the copies and rotate.