Almost every connection your software makes starts with a DNS lookup, and many outages that look like network or application failures are DNS failures in disguise: a record cached past a migration, a resolver that cannot reach an authoritative server, a DNSSEC signature that expired overnight, or a search path that turns one lookup into five.
This article builds DNS up from its parts. It explains who asks whom, what is in a DNS message, how caching and TTLs really behave, why some answers need TCP, what the encrypted transports protect and what they do not, and how DNSSEC turns answers into signed data. A short Python program sends a real query built byte by byte, and a worked migration shows how to change a record without stranding clients.
Three roles: stub, recursive, authoritative
DNS is a distributed, hierarchical database with three kinds of participant. The stub resolver is the library inside the operating system that applications call through getaddrinfo. It knows nothing about the hierarchy; it sends one query with the recursion-desired bit set to a configured resolver and waits for a final answer. The recursive resolver, run by your ISP, your cloud provider, your company or a public service, does the work: it walks the hierarchy, caches everything it learns and returns the result. Authoritative servers hold zones, the actual data for a portion of the namespace, and answer only for those zones without recursing.
The namespace is a tree read right to left. The root zone delegates com to the TLD servers, which delegate example.com to the servers its owner chose. Delegation is expressed with NS records in the parent zone. When a name server's own name lies inside the zone it serves, for example ns1.example.com for example.com, the parent must also supply its address, called glue, or the resolver could never reach it.
A resolution, step by step
Suppose the cache is empty and an application asks for www.example.com. The resolver starts from a built-in list of root server addresses, the root hints. It asks a root server; the root does not know the answer but returns a referral: the NS records for com plus glue addresses. The resolver asks a com server and receives a referral to the example.com name servers. It asks one of those and gets an authoritative answer, with the AA bit set, containing the A record and its TTL. Modern resolvers apply QNAME minimisation (RFC 9156): they reveal only the next label to each server, asking the root about com rather than the full name, which leaks less to upstream servers.
In practice the root and TLD steps are almost always cached, because their records carry TTLs of a day or more, so a typical cache miss costs one round trip to the authoritative servers. dig +trace www.example.com performs the walk itself from the root and prints each referral, which is the fastest way to see where a broken delegation stops.
What is in a DNS message
Every DNS message, query or response, has the same layout: a 12-byte header, then the question, answer, authority and additional sections. The header carries a 16-bit ID that the client picks at random and the response must echo, a flags word and four counts. The flags that matter most are QR (response), AA (authoritative answer), TC (truncated), RD (recursion desired), RA (recursion available) and the 4-bit RCODE: 0 for NOERROR, 2 for SERVFAIL, 3 for NXDOMAIN. Names are encoded as length-prefixed labels ending in a zero byte. The program below builds a query by hand, sends it over UDP, checks the ID and falls back to TCP when the response is truncated.
import random, socket, struct
QR, AA, TC, RD, RA = 0x8000, 0x0400, 0x0200, 0x0100, 0x0080
def build_query(name, qtype=1, edns_size=1232):
qid = random.getrandbits(16)
header = struct.pack("!HHHHHH", qid, RD, 1, 0, 0, 1 if edns_size else 0)
labels = name.rstrip(".").split(".")
qname = b"".join(bytes([len(l)]) + l.encode("ascii") for l in labels) + b"\x00"
question = qname + struct.pack("!HH", qtype, 1) # QTYPE, QCLASS=IN
# EDNS(0) OPT pseudo-record: root name, TYPE 41, CLASS = UDP payload size
opt = b"\x00" + struct.pack("!HHIH", 41, edns_size, 0, 0) if edns_size else b""
return qid, header + question + opt
def recv_exact(sock, n):
buf = b""
while len(buf) < n:
chunk = sock.recv(n - len(buf))
if not chunk:
raise ConnectionError("server closed the connection")
buf += chunk
return buf
def parse_header(resp, qid):
rid, flags, qd, an, ns, ar = struct.unpack("!HHHHHH", resp[:12])
if rid != qid or not flags & QR:
raise ValueError("not a response to our query")
return {"rcode": flags & 0xF, "aa": bool(flags & AA), "tc": bool(flags & TC),
"ra": bool(flags & RA), "answers": an, "authority": ns}
def query(server, name, qtype=1, timeout=2.0):
qid, msg = build_query(name, qtype)
with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as s:
s.settimeout(timeout)
s.sendto(msg, (server, 53))
resp, _ = s.recvfrom(4096)
h = parse_header(resp, qid)
if not h["tc"]:
return h
with socket.create_connection((server, 53), timeout=timeout) as s:
s.sendall(struct.pack("!H", len(msg)) + msg) # TCP: 2-byte length
n = struct.unpack("!H", recv_exact(s, 2))[0]
return parse_header(recv_exact(s, n), qid)
print(query("127.0.0.53", "example.com")) # use your configured resolver's addressReal resolvers also randomise the UDP source port, because a 16-bit ID alone is too easy to guess; that combination is the main defence against cache-poisoning by spoofed responses when DNSSEC is not in use.
Caching, TTLs and negative answers
Every record carries a TTL in seconds, set by the zone owner. A resolver may reuse a cached record until its TTL expires, and many layers cache: the recursive resolver, the operating system's stub cache, language runtimes and sometimes the application. A TTL is therefore a change-latency budget: after you edit a record, some clients may keep using the old value for up to one TTL, or longer when a layer ignores it. Java, for instance, caches lookups according to its networkaddress.cache.ttl security property, so check the effective value in your runtime rather than assuming it follows DNS.
Negative answers are cached too (RFC 2308). An NXDOMAIN or an empty answer includes the zone's SOA record, and the resolver caches the absence for the smaller of the SOA record's TTL and its MINIMUM field. The classic trap is querying a name before creating it: the resolver caches the non-existence, and the record you then add stays invisible to that resolver until the negative TTL expires.
Resolvers may clamp TTLs to their own minimum and maximum, and some implement serve-stale (RFC 8767): if authoritative servers are unreachable when a record expires, they keep answering with the expired data for a bounded time. That turns an authoritative outage into stale answers rather than failures, which is usually the right trade.
Message size, EDNS and TCP
The original protocol limited UDP responses to 512 bytes. EDNS(0) (RFC 6891) lets the client advertise a larger buffer through the OPT record seen in the code above. Large UDP responses, however, get fragmented at the IP layer, and fragments are often dropped by firewalls or abused in spoofing attacks. The DNS Flag Day 2020 coordination recommended a default EDNS buffer of 1232 bytes, small enough to avoid fragmentation on almost every path. When a response does not fit, the server sets TC and the client retries over TCP, which RFC 7766 requires every implementation to support. A firewall that blocks TCP port 53 therefore breaks large answers, and DNSSEC-signed answers are often large.
Encrypted transports: DoT, DoH and DoQ
Classic DNS is plaintext, so anyone on the path can read and alter queries. DNS over TLS (RFC 7858) runs the same messages over TLS on port 853. DNS over HTTPS (RFC 8484) carries them in HTTPS requests on port 443, which blends with web traffic and lets browsers choose their own resolver. DNS over QUIC (RFC 9250) uses QUIC streams, avoiding head-of-line blocking; QUIC architecture explains the transport.
All three protect the hop between the stub and its recursive resolver. They do not hide your queries from the resolver itself, and the hop from the resolver to authoritative servers is still mostly plain UDP and TCP. For enterprises, browser-chosen DoH can bypass internal resolvers and split-horizon names, so decide a policy and configure it rather than discovering it during an incident.
DNSSEC: signed answers and a chain of trust
DNSSEC adds signatures rather than encryption. A zone publishes its public keys in DNSKEY records and signs each record set with an RRSIG. The parent zone publishes a DS record, a hash of the child's key-signing key, and signs it with its own key. A validating resolver starts from the root's key, configured as a trust anchor, and follows DS to DNSKEY to RRSIG down to the answer. Non-existence is proven with NSEC or NSEC3 records that cover the gap where the name would be.
If validation fails, the resolver returns SERVFAIL rather than an unverified answer. That is the source of most DNSSEC outages: a DS record left in the parent after changing DNS providers, a key rollover done out of order, or signatures allowed to expire because the signer stopped. Automate signing and rollovers, monitor signature expiry, and treat DS changes at the registrar as a change with a rollback plan.
Operating DNS for real services
Host authoritative zones on at least two independent networks or providers and keep them in sync with zone transfers or provider APIs; authoritative DNS is a single point of failure for everything under the name. Large authoritative and public resolver services are spread with anycast, described in Anycast architecture. Health-checked steering and failover records belong to the authoritative layer, covered in Cloud DNS failover and traffic steering.
Inside clusters DNS is also service discovery. Kubernetes pods get a resolver configuration with several search domains and ndots:5, so a name with fewer than five dots, such as api.example.com, is first tried with each search suffix appended. One external lookup can become several queries, each multiplied by A and AAAA. Use fully qualified names ending in a dot for external hosts, or lower ndots for the pod, and run a node-local cache. Service discovery architecture compares DNS-based discovery with registries.
Worked example: moving an API to a new load balancer
The record api.example.com A has a TTL of 3600 seconds and must point at a new load balancer on Tuesday. On the preceding Monday morning, at least one old TTL before the change, lower the TTL to 60 seconds; resolvers that cached the record under the old TTL will have refetched it with the new one within the hour. On Tuesday, bring the new endpoint up, verify it directly, then change the record. Within about a minute most resolvers return the new address. Keep the old load balancer serving for at least a day, because some runtimes and devices ignore TTLs, and watch its traffic decay to zero before removing it. After a stable week, raise the TTL again to reduce query load and keep answers available through short authoritative outages. If the record is new rather than changed, do not query it before it exists, or negative caching will hide it.
Failure modes
- Lame delegation. The parent lists a name server that does not serve the zone. Some queries time out, depending on which server a resolver picks.
- Missing or stale glue. In-zone name servers renumbered without updating glue at the registrar.
- DNSSEC validation failure. Expired signatures or a mismatched DS make the whole zone SERVFAIL for validating resolvers, while non-validating ones still work, which confuses diagnosis.
- TCP blocked. Large answers fail only for some names.
- Long TTL before a migration. Clients keep the old address for hours.
- Search-path amplification. Resolver load and latency multiply inside clusters.
- Resolver outage. Every dependency fails at once; configure at least two resolvers and cache locally.
Trade-offs
Long TTLs cut latency and query cost and ride out authoritative outages, but slow every change; short TTLs do the opposite. DNSSEC prevents forged answers but adds operational steps whose failure takes the zone offline. Encrypted transports give privacy on the first hop but can bypass local policy. Serve-stale improves availability at the price of correctness during an outage. Choose per record: stable infrastructure names can carry long TTLs, failover names short ones.
What to do next
- Run dig +trace for your main domain and confirm every delegation and glue record is consistent.
- Run the Python query program against your resolver and inspect the RCODE, AA and TC bits.
- List the TTLs of your important records and decide which should be short and which long.
- Check that TCP port 53 is open from your resolvers and that EDNS buffers are 1232 bytes or less.
- If you sign zones, add monitoring for RRSIG expiry and DS consistency at the registrar.
- Inside Kubernetes, measure queries per external lookup and fix ndots or use trailing dots.
- Write a migration runbook that lowers TTLs one old TTL ahead of every change.