In most microservice platforms, DNS is the service registry. A service calls http://orders:8080 without knowing where orders runs. The platform turns that short name into an address, and the address changes every time pods are rescheduled. This works well enough that teams forget it exists, until a deploy adds 5-second stalls, a scale-up does not spread load to the new pods, or the cluster DNS server becomes the busiest service in the cluster.

This article covers DNS inside a cluster, with Kubernetes and CoreDNS as the concrete example: what records exist, how a lookup is expanded, where answers are cached, and why DNS often balances load worse than people expect. The protocol itself, including resolvers, delegation, TTLs and DNSSEC, is in the DNS architecture deep dive. Registry-based discovery in general is in service discovery architecture.

Advertisement

What the cluster publishes

Name resolution path for a pod in a Kubernetes clusterApp containergetaddrinfo / language resolver/etc/resolv.conf: search list, ndots:5UDP/TCP 53NodeLocal DNSCacheoptional, per nodemissCoreDNS podsbehind the kube-dns ServicewatchKubernetes APIServices, EndpointSlicesforward .Upstream resolverexternal namesconnect to answerClusterIP (virtual)kube-proxy / eBPF rulesper connectionPod A10.1.2.7Pod B10.1.3.4Pod C10.1.5.9Normal Service: DNS returns one ClusterIP; balancing happens per connection in the dataplane.Headless Service: DNS returns the pod IPs themselves; the client must balance.Every layer caches differently, so a name can be stale at more than one place.
Figure 1. A pod's query passes through the search-list logic in its own resolver, an optional node-local cache and CoreDNS, which answers cluster names from an API watch and forwards everything else upstream.

According to the Kubernetes DNS documentation, a normal Service gets an A or AAAA record at my-svc.my-namespace.svc.cluster-domain.example. With the usual cluster.local domain, that becomes orders.shop.svc.cluster.local. The record resolves to the Service's single virtual ClusterIP. A headless Service, one declared with clusterIP: None, gets a record with the same name, but it resolves to the IPs of all the ready pods behind it.

Named ports also get SRV records, at _port-name._port-protocol.my-svc.my-namespace.svc.cluster-domain.example. For a normal Service, the SRV answer points at the Service name and port. For a headless Service it returns one answer per pod, with per-pod hostnames. Few HTTP clients read SRV, but gRPC resolvers and some database drivers can, and it is the only DNS record type that carries the port.

CoreDNS builds these answers from a watch on Services and EndpointSlices, so a record follows the API state within moments. The kubernetes plugin's default TTL is 5 seconds, configurable from 0 up to 3600. Per-pod A records based on IP addresses are disabled by default in the plugin. Many installed Corefiles turn them on with pods insecure for compatibility, so check what your cluster actually runs.

The search list and ndots:5

The kubelet writes the pod's /etc/resolv.conf. With the default ClusterFirst DNS policy it looks like this:

nameserver 10.96.0.10
search shop.svc.cluster.local svc.cluster.local cluster.local
options ndots:5

ndots:5 means that a name containing fewer than five dots is tried with each search suffix before it is tried as an absolute name. That is what lets orders resolve to orders.shop.svc.cluster.local on the first attempt. It also means that every external name with fewer than five dots walks the whole search list first.

Worked example. A pod in namespace shop calls api.payments.example.com. That name has three dots, which is fewer than five, so a glibc-style resolver tries:

api.payments.example.com.shop.svc.cluster.local.   -> NXDOMAIN
api.payments.example.com.svc.cluster.local.        -> NXDOMAIN
api.payments.example.com.cluster.local.            -> NXDOMAIN
api.payments.example.com.                          -> answer

Many resolvers ask for A and AAAA in parallel, so that is eight queries, six of them wasted, for every lookup that is not cached. Multiply by the number of pods and by languages that do not cache, and CoreDNS ends up spending most of its time on NXDOMAIN answers for external names. A short script shows the expansion for any name:

def expansions(name, search, ndots=5):
    if name.endswith("."):
        return [name]                          # absolute: no search list
    absolute = name + "."
    suffixed = [f"{name}.{s}." for s in search]
    return [absolute] + suffixed if name.count(".") >= ndots else suffixed + [absolute]

search = ["shop.svc.cluster.local", "svc.cluster.local", "cluster.local"]
for n in ["orders", "orders.billing", "api.payments.example.com", "api.payments.example.com."]:
    q = expansions(n, search)
    print(f"{n:28s} {len(q)} names, {2 * len(q)} queries worst case (A + AAAA)")

There are three ways to fix it. Write external names fully qualified with a trailing dot (api.payments.example.com.). This is the cheapest option, but some TLS and HTTP libraries mishandle the dot in the Host header or SNI, so test it. Alternatively, lower ndots for pods that mostly call external services:

spec:
  dnsPolicy: ClusterFirst
  dnsConfig:
    options:
      - name: ndots
        value: "2"

With ndots:2, orders and orders.billing (the service-dot-namespace form) still use the search list, and names with two or more dots try the absolute name first, falling back to the search list only on NXDOMAIN. The third fix is node-local caching, so the wasted queries never leave the node. That is covered next.

Advertisement

Caching layers, and which TTL actually applies

A cluster name can be cached in up to four places, and each follows different rules.

LayerWhat it cachesNotes
CoreDNS cache pluginPositive and negative answersDefault Corefiles commonly set cache 30 (seconds); cluster records still carry the plugin's 5-second TTL
NodeLocal DNSCachePer-node cache in front of CoreDNSListens on a link-local address (the docs use 169.254.20.10 as the example); can upgrade connections to CoreDNS to TCP
Language runtimeVariesGo's resolver and glibc getaddrinfo do not cache by themselves; OpenJDK caches positive answers for an implementation-specific period (30 seconds is common when no security manager is installed) controlled by networkaddress.cache.ttl; Node.js dns.lookup does not cache
Connection poolThe connection itselfNo TTL at all: a pooled connection to a pod lives until it is closed

The last row matters most. DNS answers only matter when a client opens a new connection. A service that holds an HTTP/2 or gRPC connection for hours does not look at DNS again during that time, whatever the TTL says.

Why DNS balances worse than you expect

With a normal Service, DNS returns one ClusterIP, and the dataplane (kube-proxy iptables or IPVS rules, or an eBPF replacement) picks a backend pod per connection. HTTP/1.1 clients with many short connections spread well. A gRPC client opens one HTTP/2 connection and sends every request over it, so all of its traffic goes to one pod. Scale orders from 3 to 10 pods and existing clients keep using the original 3 until their connections break.

The usual fix is to make the client do the balancing. Point it at a headless Service so DNS returns every pod IP, and enable a round-robin policy. In grpc-go:

conn, err := grpc.NewClient(
    "dns:///orders-headless.shop.svc.cluster.local:50051",
    grpc.WithTransportCredentials(creds),
    grpc.WithDefaultServiceConfig(`{"loadBalancingConfig":[{"round_robin":{}}]}`),
)

This brings a second problem. gRPC's DNS resolver does not poll on the TTL. It re-resolves mainly when connections fail or close, so newly added pods are not discovered until something triggers a refresh. Set a maximum connection age on the server (in grpc-go, keepalive.ServerParameters{MaxConnectionAge: ...}) so clients reconnect and re-resolve regularly. Minutes is a typical scale. A service mesh moves balancing into a sidecar or node proxy that watches endpoints directly and does not depend on DNS. That is the main reason meshes balance gRPC well, as the service mesh architecture guide explains.

Headless Services have their own failure modes. A DNS answer lists every ready pod. A client that caches it keeps sending to pods that have since terminated, and gets connection refused or timeouts until it re-resolves. Clients must treat a failed connection as a reason to re-resolve, not just to retry the same address.

Worked example: the scale-up that did not help

Consider a checkout service that calls orders over gRPC through a normal ClusterIP Service. Checkout runs 4 replicas, each with one long-lived channel. Orders runs 3 pods. At peak, orders CPU reaches 90 percent, and the autoscaler adds 5 more pods. Nothing changes. The 4 checkout channels were spread over the original 3 orders pods when they were opened, and with no reconnection none of them moves. The 5 new pods sit idle and are scaled back down later.

Going through the layers explains why. DNS answered orders correctly with one ClusterIP. The dataplane chose a pod per connection, correctly. The problem is that there were only 4 connections, so there were only 4 balancing decisions. The fix has three parts. Switch checkout to a headless Service with round_robin, so each channel connects to every pod. Set MaxConnectionAge on orders, so channels reconnect and re-resolve after a scale event. Then confirm the change by watching per-pod request rates, not pod count, during the next load test. The lesson generalises: whenever load does not follow replicas, count connections before you blame DNS.

Failure modes seen in production

  • Five-second stalls. glibc sends the A and AAAA queries from the same UDP socket at the same moment. On some kernels, a race in conntrack insertion drops one of them, and the resolver waits for its default 5-second timeout before retrying. The symptom is a latency bump at exactly 5 seconds. Mitigations include options single-request-reopen or use-vc in dnsConfig, NodeLocal DNSCache (which removes conntrack from the path), and newer kernels. Which ones help depends on your kernel and libc versions, so measure.
  • CoreDNS saturation. Too few replicas for the query rate, often caused by ndots amplification, show up as SERVFAIL and timeouts across the cluster. Scale CoreDNS with node or core count (the cluster-proportional autoscaler is a common choice), and fix amplification before adding replicas.
  • Negative caching hides new Services. A client that looked up a Service before it existed may keep an NXDOMAIN cached. Deploy dependencies first, and keep negative TTLs short.
  • Upstream outages leak in. If the forward upstream fails, every external lookup that is not cached fails as well. Check the forward targets, and decide whether CoreDNS should serve stale entries during an outage. The cloud DNS failover guide covers the external side.
  • Names baked into config outlive the cluster layout. A config file that says orders.shop.svc.cluster.local breaks if the service moves namespace or the cluster domain is not cluster.local. Inject service addresses from the deployment environment rather than hard-coding the domain, and test a namespace rename in staging.
  • Alpine images behave differently. musl's resolver handles search lists and parallel queries differently from glibc. Test the image you ship, not a debug image built on another base.

Operating it

CoreDNS exposes Prometheus metrics through its prometheus plugin, commonly on port 9153. Watch request rate (coredns_dns_requests_total), responses by rcode (coredns_dns_responses_total), cache hit ratio and request duration. An NXDOMAIN share above half usually means search-list amplification. A rising SERVFAIL share means trouble with an upstream or the API watch.

When debugging, run the client's own resolver from inside the affected pod, not from your laptop:

kubectl exec -n shop deploy/checkout -- cat /etc/resolv.conf
kubectl run dnsdebug -n shop --rm -it --image=<image with dig> -- \
    dig +search +showsearch api.payments.example.com
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=100

The +showsearch option prints every expanded query, which makes amplification visible in a single command. Enable the CoreDNS log plugin only briefly, because under load it is expensive.

Trade-offs: DNS, registry or mesh

ApproachStrengthWeakness
ClusterIP + DNSZero client code, any languageBalances per connection; pinned under HTTP/2
Headless + client-side balancingPer-request balancing, no extra hopClient must re-resolve; stale pod lists
Registry API (watch endpoints)Push updates, rich metadataClient library per language
Service mesh proxyBalancing, retries and mTLS outside the appExtra hop, more moving parts to operate

What to do next

  1. Print /etc/resolv.conf from three representative pods and confirm the search list and ndots.
  2. Use dig +search +showsearch on your top external dependency and count the queries.
  3. Lower ndots or use trailing-dot FQDNs for external-heavy services, and measure the drop in CoreDNS NXDOMAIN rate.
  4. Find every gRPC or HTTP/2 client and check how it balances: headless with round_robin, a mesh, or pinned.
  5. Set a server max connection age so clients re-resolve, and check that new pods take traffic within minutes.
  6. Add CoreDNS dashboards for rcode mix, cache hit ratio and latency, and alert on SERVFAIL and 5-second latency spikes.
Key takeaway: Inside a cluster, DNS is the service registry. Kubernetes publishes an A or AAAA record per Service, pod IPs for headless Services, and SRV records for named ports. CoreDNS answers from an API watch with a 5-second TTL by default. ndots:5 and the search list make every short external name cost several wasted queries, so qualify names, lower ndots, or cache on the node. DNS only affects new connections. Long-lived HTTP/2 and gRPC connections pin to pods unless the client balances over a headless Service and reconnects regularly, or a mesh does that work. Watch the rcode mix and look for 5-second latency spikes.