Cloud DNS is Google Cloud's managed authoritative DNS service and the resolver behind every VPC network. Most teams set it up once, with a public zone for the website, a private zone for internal names and maybe a forwarding zone to on-premises, and then forget it until an outage. The outages are rarely about the zone design. They come from operations: a record changed by hand with a one-day TTL still cached, a Kubernetes cluster whose service names nobody outside can resolve, an override that nobody remembers adding, or a resolution failure no one can see because logging was never turned on.
This article is about running Cloud DNS. It assumes you know the zone types and resolution order, which are covered in GCP Cloud DNS, in depth. It covers managing records as code, planning TTLs for cutovers, response policies for overriding answers per network, Cloud DNS for GKE and its scopes, query logging and metrics, and a troubleshooting playbook. Commands and flags are the ones in Google's documentation at the time of writing. Check the current docs before automating against them.
The resolution model in one paragraph
A short recap, so the operations make sense. Public zones are authoritative on the internet once your registrar delegates to the zone's Google name servers. Private zones answer only for the VPC networks they are attached to. Forwarding zones send matching queries to resolvers you name, usually on-premises. Peering zones let one network resolve using another network's DNS configuration. An inbound server policy gives on-premises clients an address in your VPC to query. Every VM sends queries to the metadata server at 169.254.169.254, and the VPC resolver decides which of these answers.
Records as code
Treat DNS records like application code. Every change goes through review, every zone has an owner, and drift from hand edits is detected and reverted. The Cloud DNS API applies a change as a set of additions and deletions, so replacing a record's data is one atomic change. gcloud is fine for one-off work and break-glass use.
# One-off: create, then update, an A record (note the trailing dot on the name)
gcloud dns record-sets create api.example.com. --zone=example-public \
--type=A --ttl=300 --rrdatas=203.0.113.10
gcloud dns record-sets update api.example.com. --zone=example-public \
--type=A --ttl=60 --rrdatas=203.0.113.20For everything else, use Terraform or another declarative tool, so the repository is the source of truth.
resource "google_dns_record_set" "api" {
name = "api.${google_dns_managed_zone.public.dns_name}"
managed_zone = google_dns_managed_zone.public.name
type = "A"
ttl = var.api_ttl # lowered before a cutover, raised after
rrdatas = var.api_ips
}Three conventions make this work at scale. First, one state file per zone, so a bad plan cannot touch every zone at once. Second, a CI check that fails if a plan deletes more than a handful of records, because a misread import can make Terraform want to delete a zone's contents. Third, a scheduled drift job that runs a plan and alerts on any diff, so console edits do not survive quietly.
Watch for split-horizon surprises. If a private zone attached to a network covers example.com, VMs in that network get their answers from the private zone, not from the public zone. A record that exists only in the public zone does not resolve from inside, which looks like a broken record when it is really a missing one. Either keep the private zone narrow, such as internal.example.com, or mirror the public records the network needs, and generate both from one source in code so they cannot drift.
TTLs and cutover planning
A TTL is a promise that resolvers may keep using an answer for that many seconds. You cannot recall a cached answer. The only control is to shorten the TTL before you need the change to take effect. Plan cutovers as a timeline. Suppose api.example.com has a TTL of 86,400 seconds and you want to move it on Friday. Lower the TTL to 60 on Wednesday or earlier, so every cache holding the old one-day value has expired by Friday. Make the change on Friday, and the world converges within about a minute, give or take resolvers that ignore TTLs. After a day of stability, raise the TTL again to cut query volume and lookup latency.
Negative answers are cached too. When a name does not exist, resolvers cache the NXDOMAIN for a period taken from the zone's SOA record. If clients queried a name before you created it, some will keep failing until that negative TTL expires. Create new names ahead of the traffic that needs them, and check the SOA values rather than assuming them.
Response policies
A response policy is a set of rules, attached to VPC networks, that changes what the resolver returns for selected names. A rule either supplies local data, records returned instead of whatever private zones, Google Cloud internal DNS or the public internet would have said, or applies a behavior. The documented behavior is bypassResponsePolicy, which skips a less specific rule and lets resolution continue as if the policy had not matched. Only one response policy can be attached to each network, and rules select by DNS name, including wildcards.
The classic use is pointing Google APIs at a private endpoint. A wildcard rule sends every googleapis.com name to the restricted.googleapis.com VIP, so traffic stays on the path that VPC Service Controls protects. Bypass rules then exempt the few names that must keep their normal answers.
gcloud dns response-policies create corp-overrides --networks=prod-vpc \
--description="API routing and staged migrations"
gcloud dns response-policies rules create googleapis-restricted \
--response-policy=corp-overrides --dns-name="*.googleapis.com." \
--local-data=name="*.googleapis.com.",type="CNAME",ttl=300,rrdatas="restricted.googleapis.com."
gcloud dns response-policies rules create restricted-vip \
--response-policy=corp-overrides --dns-name="restricted.googleapis.com." \
--local-data=name="restricted.googleapis.com.",type="A",ttl=300,rrdatas="199.36.153.4"
gcloud dns response-policies rules create keep-normal-answer \
--response-policy=corp-overrides --dns-name="special.googleapis.com." \
--behavior=bypassResponsePolicyThe wildcard also matches restricted.googleapis.com itself, and response policy data overrides private zones, so the more specific restricted-vip rule is required, or the CNAME would point at itself. Google documents four addresses for that name, 199.36.153.4 to 199.36.153.7; only one is shown here, so add the rest. special.googleapis.com is a placeholder for a name to exempt. Response policies are also useful for staged migrations: override one name for one network to test a new backend before changing the authoritative record. Because they override everything, they are dangerous when forgotten. Give every rule an owner and an expiry in its description, and review them like firewall rules.
Cloud DNS for GKE and its scopes
By default a GKE Standard cluster runs kube-dns pods inside the cluster. Cloud DNS for GKE moves service-name resolution into Cloud DNS, so the kube-dns pods and their scaling problems are no longer needed. Pods on Standard clusters then use the metadata server, 169.254.169.254, as their nameserver, and NodeLocal DNSCache is used where it is enabled, as it is by default on Autopilot. New Autopilot clusters from version 1.25.9-gke.400 default to Cloud DNS with cluster scope. There are three scopes.
| Scope | Who can resolve service names | Flags at creation | Notes |
|---|---|---|---|
| Cluster | Only pods in the cluster | --cluster-dns=clouddns --cluster-dns-scope=cluster | Same behaviour as kube-dns; the default |
| VPC | Any client in the VPC, under a custom domain | --cluster-dns=clouddns --cluster-dns-scope=vpc --cluster-dns-domain=DOMAIN | Not supported on Autopilot; lets non-GKE clients find headless services |
| Additive VPC | Cluster scope plus a VPC-visible zone | --cluster-dns=clouddns --cluster-dns-scope=cluster --additive-vpc-scope-dns-domain=DOMAIN | Autopilot: only at cluster creation |
Two operational facts catch teams out. The documentation states that the DNS scope cannot be changed after it is set, so choose scope at cluster creation and treat a change as a migration to a new cluster. And enabling Cloud DNS on an existing cluster does not remove kube-dns. It keeps running until you scale it down yourself, which is the moment to watch resolution errors closely. Pick a unique domain per cluster for VPC or additive scope, so two clusters never publish the same service names into one network.
Query logging and metrics
Without logging, DNS failures show up as application timeouts with no explanation. For private zones and VPC resolution, enable logging on a DNS server policy attached to the network. For public zones, enable it on the zone. Entries land in Cloud Logging under the dns_query resource type. Google's documentation estimates roughly 5 MB of log data per 10,000 queries, and while Cloud DNS does not charge for logging, Cloud Logging storage does cost money. Use exclusion filters or shorter retention for noisy, healthy traffic.
gcloud dns policies create prod-logging --networks=prod-vpc --enable-logging
gcloud dns managed-zones update example-public --log-dns-queries
# Logs Explorer: failing lookups from VMs in the last hour
resource.type="dns_query"
(jsonPayload.responseCode="NXDOMAIN" OR jsonPayload.responseCode="SERVFAIL")Useful fields include queryName, queryType, responseCode, sourceIP, vmInstanceName, source_type (values such as gce-vm, inbound-forwarding and peering-zone) and egressError, which records failures reaching a forwarding target. For dashboards and alerts, the dns.googleapis.com/query/response_count metric counts responses by response code. A rising SERVFAIL rate usually means a forwarding target or DNSSEC problem. A rising NXDOMAIN rate usually means a bad deploy or a search-path storm.
A troubleshooting playbook
When something does not resolve, work through the layers in order instead of guessing.
- Reproduce from the same network: on a VM in the affected VPC, run
dig @169.254.169.254 name.example.internaland note the status and answer. - Check response policies attached to the network first, because a matching rule overrides everything below it.
- Check which private, forwarding or peering zone is the most specific match for the name, and that it is attached to this network.
- For forwarding zones, look for egressError in the logs and confirm that firewall rules and routes let Google's forwarding source range reach the target resolvers.
- For public names, query the zone's own name servers directly to separate authoritative data from cached data, then compare with a public resolver.
- Check the TTL on the answer you received. If the record is correct at the source but the answer is wrong, the problem is caching. Wait or flush where you control the cache.
Worked example: moving one internal name
Worked example: a team is moving orders.corp.example from an on-premises load balancer to an internal load balancer in Google Cloud. Today, a forwarding zone sends corp.example to on-premises resolvers. A week before the move they lower the on-premises record's TTL to 60 seconds. On migration day, they first add a response policy rule on a test network only, giving orders.corp.example the new internal IP, and run the smoke tests from that network. Production is unaffected because the rule is attached only to the test network.
Once the tests pass they make the real change. They create a private zone for orders.corp.example in Terraform, attached to the production VPC. Because it is more specific than the corp.example forwarding zone, it wins for that name. Within a minute, query logs show orders.corp.example answered by the private zone with NOERROR. They delete the test rule the same day, and raise the TTL after a week.
Failure modes
- Forgotten response policy rule. An override added for a test outlives the test and silently pins a name to a dead IP.
- Cutover without a TTL drop. The record changes, but clients keep the old answer for up to the previous TTL.
- Overlapping cluster domains. Two GKE clusters in VPC scope publish the same names into one network.
- kube-dns scaled down too early. Pods still configured for kube-dns lose resolution mid-migration.
- Forwarding target unreachable. A firewall change blocks Google's forwarding source range, producing SERVFAIL everywhere that zone applies.
- No logging when it matters. Logging enabled after an incident cannot explain the incident.
Trade-offs
Short TTLs make changes fast and failover responsive but raise query volume and add lookup latency. Response policies are precise and immediate but are invisible to anyone reading only the zones. VPC-scope GKE DNS makes services discoverable to VMs but ties the scope decision to cluster creation. Full query logging gives complete visibility at a storage cost, so many teams log everything in production networks and use exclusion filters for known-good, high-volume names. Declarative management is slower for a single urgent fix than a console edit, so keep a documented break-glass path with gcloud and require the change to be back-ported to code the same day.
What to do next
- Import every zone into Terraform, one state per zone, and add a drift check.
- List current TTLs and write a cutover runbook that lowers them days in advance.
- Inventory response policy rules and give each an owner and an expiry.
- Choose a GKE DNS scope per cluster before creating it, with a unique domain for VPC or additive scope.
- Enable query logging on production networks and public zones, and alert on SERVFAIL and NXDOMAIN rates.
- Rehearse the troubleshooting playbook once. Further reading: GCP VPC, Private Google Access and Private Service Connect, Shared VPC and DNS in microservices.