Cloud DNS is two products under one name. For the internet it is an authoritative DNS service that answers for your public domains from Google's anycast name servers. Inside Google Cloud it is the resolver behind every VPC network: when a VM asks the metadata server at 169.254.169.254 to resolve a name, Cloud DNS decides the answer using private zones, policies and forwarding rules you configure. Most production incidents involving Cloud DNS come from mixing up these two roles, or from not knowing in which order the private side checks its sources.

This article explains both sides from first principles, with the resolution order at the centre. It covers the zone types, public delegation and a safe migration, DNSSEC, hybrid DNS with on-premises networks, hub-and-spoke designs with peering zones, and routing policies. It ends with a worked design for a company with a data centre and many projects, and a checklist. Commands use the gcloud dns CLI. Details were checked against the Cloud DNS documentation in October 2026, and you should check limits and preview features again before relying on them.

Advertisement

The zone types and what each one is for

A managed zone holds the records for one DNS name suffix, such as example.com.. Its visibility and type decide who can query it and where the answers come from.

Zone typeWho can query itWhere answers come fromTypical use
PublicAnyone on the internetRecords you store in the zoneYour company's domains
PrivateOnly the VPC networks you authoriseRecords you store in the zoneInternal names such as db.corp.internal.
ForwardingAuthorised VPC networksTarget name servers you list, for example on premisesResolving the data centre's domains from Google Cloud
PeeringAuthorised VPC networks (the consumer)The resolution order of another VPC network (the producer)Hub-and-spoke: spokes resolve through the hub
Response policyNetworks or GKE clusters the policy is bound toRules that return local data or bypassOverriding names, blocking domains, private API endpoints
Managed reverse lookupAuthorised VPC networksCompute Engine internal data for PTR queriesReverse lookups of private addresses
Service DirectoryAuthorised VPC networksA Service Directory namespaceService discovery by DNS

Private, forwarding and peering zones are all scoped to a list of VPC networks. A private zone that is not bound to a network does nothing for that network, and a VM in an unbound network falls through to public DNS for the same name. That fall-through is behind many "works in one project, fails in another" tickets.

How a VM's query is resolved

When a VM queries the metadata server, Cloud DNS works through five steps in order. The first step that produces an answer, including an authoritative NXDOMAIN, ends resolution.

  1. Alternative name servers. If an outbound server policy applies to the network, queries go to the alternative name servers it lists. Cloud DNS ranks them by success rate and round-trip time. The remaining steps do not run for this network.
  2. Response policies. Rules are matched by longest suffix. A rule either serves local data, which completes resolution, or bypasses the policy so resolution continues.
  3. Private zones. The most specific authorised zone wins. A private zone returns its records, or NXDOMAIN if the name is not in it. It does not fall through to public DNS for names inside its suffix. A forwarding zone sends the query to its targets. A peering zone restarts resolution in the producer network.
  4. Compute Engine internal DNS. Names of VMs resolve to their internal addresses.
  5. Public DNS. Anything left is resolved on the internet.
VM query169.254.169.2541. Alternative name serversoutbound server policy: if set, used first2. Response policieslocal data or bypass, longest suffix3. Private zonesauthoritative, forwarding, peering4. Compute Engine internal DNSVM names5. Public DNSinternet resolutionOn-prem DNSforwarding targetfrom 35.199.192.0/19Hub VPC zonesvia peering zoneOn-prem clientInbound entry pointone IP per subnetVPN / InterconnectThe first step that produces an answer ends resolution; later steps never run.
The VPC resolution order. An outbound server policy with alternative name servers takes over the network's resolution before any private zone is consulted. Queries forwarded to private targets come from 35.199.192.0/19. On-premises clients reach Cloud DNS through inbound entry points.

Two consequences follow. First, a private zone for example.com. hides the public example.com from that network completely. If you create it to override one name, you must copy every other name you still need, or use a response policy rule instead. Second, setting alternative name servers means on-premises DNS answers everything, including names in your private zones, unless those servers forward back to Google Cloud.

Advertisement

Public zones: delegation and migration

Creating a public zone allocates a set of name servers such as ns-cloud-a1.googledomains.com. through ns-cloud-a4. The zone is not live until the registrar's NS records for the domain point at those servers. Adding records is a separate step:

gcloud dns managed-zones create example-public \
  --dns-name="example.com." --visibility=public \
  --description="Public zone for example.com"

gcloud dns managed-zones describe example-public --format="value(nameServers)"

gcloud dns record-sets create www.example.com. --zone=example-public \
  --type=A --ttl=300 --rrdatas=203.0.113.10

To migrate from another provider without an outage: import or recreate every record, and compare answers from the new servers with the old ones by querying each name directly against a Cloud DNS name server. Lower the NS TTL at the old provider if it allows. Then change the delegation at the registrar and keep the old provider serving identical records until the parent zone's NS TTL has passed, which is often a day or two. If the old zone was DNSSEC-signed, disable signing there and remove the DS record at the registrar before you switch, or plan a proper key rollover. A DS record that validating resolvers cannot match to your new keys makes the whole domain fail for them.

DNSSEC in practice

Cloud DNS can sign public zones for you. Enable it on the zone, then publish the DS record for its key-signing key at your registrar. The chain of trust is only complete when the DS record is in the parent zone, and until then nothing validates. The dangerous direction is turning it off. Remove the DS record at the registrar first, wait for the DS TTL to expire at resolvers, and only then disable signing. Doing it in the other order makes validating resolvers, which include most large public resolvers, return SERVFAIL for every name in the zone.

DNSSEC also limits routing policies. With DNSSEC enabled, a health-checked routing policy item can hold only one IP address. Plan record layouts with this in mind before you sign a zone that already uses routing policies.

Hybrid DNS: inbound and outbound

Most enterprises need names to resolve both ways between Google Cloud and a data centre. Cloud DNS handles each direction with a different feature.

Google Cloud to on-premises uses a forwarding zone, for example for corp.example.com., whose targets are the data centre's DNS servers. Cloud DNS sends these queries from the range 35.199.192.0/19, which is only reachable from inside Google Cloud and from networks connected to a VPC. The on-premises network therefore needs a route back to that range through the Cloud VPN tunnel or Interconnect attachment of the same VPC that sent the query. With dynamic routing, add 35.199.192.0/19 as a custom route advertisement on the Cloud Router's BGP session. On-premises firewalls must allow TCP and UDP port 53 from that range. If the route is missing, queries go out but replies never arrive, and forwarding looks like a timeout rather than an error.

On-premises to Google Cloud uses an inbound server policy. Enabling inbound forwarding allocates one internal address in every subnet's primary range. On-premises DNS servers then conditionally forward your private zone suffixes to one of these entry points over VPN or Interconnect. The entry points are not reachable across VPC Network Peering or Network Connectivity Center, so on-premises must forward to entry points in the network that owns the resolution, usually the hub.

# On-prem -> Cloud: inbound entry points in the hub, with query logging on.
gcloud dns policies create hub-inbound --networks=hub-vpc \
  --enable-inbound-forwarding --enable-logging \
  --description="Inbound DNS from the data centre"

# Cloud -> on-prem: forward corp.example.com to the data centre resolvers.
gcloud dns managed-zones create corp-forward --visibility=private \
  --dns-name="corp.example.com." --networks=hub-vpc \
  --forwarding-targets=10.10.0.53,10.10.1.53 \
  --description="Forward corp names on premises"

Hub and spoke with peering zones

With dozens of projects, you do not want to bind every private and forwarding zone to every spoke network. Put all zones in a hub VPC. In each spoke, create a peering zone for each suffix, or for the root . if the hub should resolve everything, targeting the hub network. A query from a spoke then restarts resolution in the hub, which finds the private zone or forwards on premises from there. The forwarded query leaves from the hub, so the 35.199.192.0/19 return route only needs to exist on the hub's hybrid links. DNS peering is separate from VPC Network Peering and does not need it.

Routing policies and health checks

A record set can carry a routing policy instead of a fixed answer. Three types exist. Weighted round robin (WRR) gives each item a weight from 0 to 1000 and returns items in proportion. Geolocation (GEO) returns the item for the Google Cloud region closest to the query's source. Failover returns a primary target while it is healthy and a geolocation backup otherwise, optionally sending a trickle ratio of traffic to the backup so it stays warm. WRR and geolocation cannot be combined in one policy.

gcloud dns record-sets create api.example.com. --zone=example-public \
  --type=A --ttl=30 --routing-policy-type=WRR \
  --routing-policy-data="900=203.0.113.10;100=203.0.113.20"

Health checking depends on the zone. In private zones it covers internal load balancers: internal Application Load Balancers, internal passthrough Network Load Balancers, and internal proxy Network Load Balancers, which were in preview at the time of writing. In public zones it covers external endpoints given as public IP address and port, with check intervals from 30 to 300 seconds. Only TCP, HTTP and HTTPS checks are supported. Remember that DNS failover is only as fast as the record TTL plus the detection time plus client caching. For fast regional failover of HTTP traffic, a global load balancer with a single anycast address is usually better. DNS policies suit non-HTTP protocols, internal load balancers and deliberate traffic splits.

Worked example: one company, one data centre, forty projects

A company has an on-premises Active Directory domain corp.example.com, forty application projects, each with its own spoke VPC, and a hub VPC that holds the Interconnect. The requirements are: every workload resolves both on-premises names and internal Google Cloud names, the data centre resolves gcp.example.internal, and workloads reach Google APIs through private addresses.

  1. In the hub, create a private zone gcp.example.internal. for internal service names and a forwarding zone corp.example.com. targeting the domain controllers. Bind both to the hub VPC.
  2. Create a private zone googleapis.com. in the hub that maps names to the private Google API range chosen for private access, as described in Private Google Access and Private Service Connect. Remember this zone hides all public googleapis.com records, so it needs the wildcard CNAME.
  3. In each spoke, create peering zones for corp.example.com., gcp.example.internal. and googleapis.com. targeting the hub. Put this in the project factory's Terraform so every new project gets it.
  4. Create the inbound server policy on the hub. On-premises DNS conditionally forwards gcp.example.internal to the hub entry points.
  5. Advertise 35.199.192.0/19 from the hub's Cloud Router, open port 53 from it on the data centre firewalls, and enable query logging on the hub.

Test from a spoke VM with dig for each suffix, and from on premises against an entry point. Then test the failure path: shut one forwarding target and check that answers still arrive.

Failure modes

  • Shadowed public names. A private zone for a public suffix returns NXDOMAIN for every name you did not copy into it.
  • Silent forwarding timeouts. No return route for 35.199.192.0/19 on the hybrid link, or a firewall blocking port 53.
  • Resolution loops. On-premises forwards a suffix to Cloud DNS, and a Cloud DNS forwarding zone sends the same suffix back on premises.
  • Unbound networks. A new spoke without peering zones resolves internal names publicly or not at all.
  • DNSSEC lockout. Signing disabled while a DS record is still published, or a provider migration without a key rollover.
  • Over-trusting DNS failover. Clients and resolvers that cache beyond the TTL keep sending traffic to the failed endpoint.

Operations and trade-offs

Manage zones and records in Terraform (google_dns_managed_zone, google_dns_record_set, google_dns_policy) rather than in the console. Review DNS changes like code, because a wrong record is an outage. Enable query logging on server policies in the networks you care about, knowing it adds log volume and cost. Keep TTLs moderate: long TTLs survive provider trouble better, and short TTLs make changes and failover faster. Centralising in a hub gives one place to reason about names, at the cost of a shared dependency that needs change control. Compare the model with Amazon Route 53, and see GCP load balancing and GCP VPC networks for the layers around it.

What to do next

  1. List every managed zone in your organisation with its type and bound networks, and find networks with no binding.
  2. Draw the five-step resolution order for your main network and mark which step answers each important suffix.
  3. Check that the hybrid links advertise 35.199.192.0/19 and that on-premises firewalls allow port 53 from it.
  4. Move spoke DNS to peering zones targeting a hub, and add them to your project factory.
  5. For each public zone, record whether DNSSEC is on and where the DS record lives, and write the disable-in-the-right-order runbook.
  6. Turn on query logging in the hub and build a dashboard of NXDOMAIN and SERVFAIL rates by suffix.
Key takeaway: Cloud DNS is an authoritative service for public zones and the resolver for every VPC network. Inside a network it checks alternative name servers, then response policies, then private, forwarding and peering zones, then internal VM names, then public DNS, and the first step that answers wins. Build hybrid DNS with forwarding zones out, inbound entry points in, and a return route for 35.199.192.0/19. Centralise zones in a hub reached through peering zones, sequence DNSSEC changes carefully, and treat DNS routing policies as a tool for specific cases, not a replacement for load balancing.