Turning on HTTP/3 is a configuration change. Running it in production is an experiment: you are adding a second transport over UDP, alongside the TCP one you cannot remove, and asking clients on networks you do not control to choose between them. Done carelessly, you get a feature that is "on" for an unknown fraction of users, with no evidence it helps and no quick way to turn it off.

This article is the runbook for doing it deliberately. The protocol itself, packets, the combined transport and TLS handshake, QPACK and connection migration, is covered in HTTP/3 and QUIC architecture and the transport in QUIC architecture, including server configuration, socket buffers and offloads. Here the focus is everything around the configuration: prerequisites, staged advertisement, measurement that survives selection bias, 0-RTT policy, abuse handling across a fleet, deploys, and the kill switch.

Advertisement

What changes when you turn it on

Clients almost never start with HTTP/3. They connect over TCP, learn that HTTP/3 is available, and on later connections race QUIC against TCP, using whichever wins. So enabling HTTP/3 splits your traffic into three populations: clients that never try, because they do not support it or have not seen the advertisement; clients that try and fail, because a network drops or throttles UDP; and clients that succeed. Your goals are to know the size of each, to make failure cheap, and to measure whether success is actually faster.

Two properties make this safe in principle. Browsers keep TCP as a fallback and, in the spirit of Happy Eyeballs, do not wait long for a QUIC path that is not working. Browsers also remember failures and avoid HTTP/3 to that origin for a while. Neither property helps if your own infrastructure breaks the UDP path for a subset of servers, which is the failure you are most likely to cause.

Architecture of a guarded rollout

Rolling out HTTP/3: discovery, a guarded UDP path, TCP fallback, and telemetry on bothClientbrowser or appDNS HTTPS recordalpn=h2,h3TCP 443: h2Alt-Svc: h3 (cohort)UDP 443edge LB / CDNrace QUICQUICterminatorOrigin over h1/h2unchanged backendAccess logs$http3, request_timeRUMnextHopProtocolTransport metricsUDP drops, resets
Discovery comes from DNS or from an Alt-Svc header sent to a chosen cohort over TCP. The UDP path terminates at the edge, the origin is unchanged, and telemetry is collected for both protocols so cohorts can be compared.

Terminate QUIC at the edge and keep the origin on HTTP/1.1 or HTTP/2; the benefits of HTTP/3 are on the lossy, high-latency client link, not inside your data center. If you use a CDN, this is usually a toggle on its side, and most of this runbook still applies to measurement and rollback. See how CDNs work for where termination sits.

Advertisement

Prerequisites

  • UDP 443 open end to end: cloud security groups, network firewalls and host firewalls all default to thinking about TCP. Test from outside, not from inside the VPC.
  • A load balancer that forwards UDP with flow affinity at least as good as the client's NAT binding. Routing on server identifiers embedded in connection IDs, as the IETF QUIC-LB draft describes, is better; four-tuple hashing breaks on migration and NAT rebinding.
  • TLS 1.3 on the QUIC listener, with TLS 1.2 still offered on TCP for older clients.
  • UDP DDoS capacity at the edge. Many mitigations rate-limit UDP 443 aggressively by default because it historically carried little legitimate traffic.
  • Logging that records the protocol per request, set up before the first advertisement, so there is a baseline.

Staged advertisement and rollback

HTTP/3 has two discovery mechanisms, and their rollback behaviour differs. An Alt-Svc response header (RFC 7838) tells a client that the same origin is reachable over h3; its ma parameter is how many seconds the client may remember that, and the default if omitted is 24 hours. An HTTPS DNS record (RFC 9460) lists h3 in its ALPN values so clients can try QUIC on the very first connection; it is remembered for the record's DNS TTL. Details of the record type are in DNS architecture.

Start with Alt-Svc only, sent to a percentage of clients, with a short ma. The cohort you do not advertise to is your holdout. Because advertisement is cached by clients, keep the max-age short, five to ten minutes, during rollout: rollback then converges in minutes, and you can lengthen it once the path is boring. Add the HTTPS record last, with a short TTL, when you are confident.

# nginx: advertise HTTP/3 to 5% of clients, keyed by address so a client stays in its cohort.
split_clients "${remote_addr}h3" $h3_cohort {
    5%   "on";
    *    "off";
}

map $h3_cohort $alt_svc {
    "on"    'h3=":443"; ma=600';
    default "";
}

server {
    # ... listen directives for TCP and QUIC as in your existing config ...
    add_header Alt-Svc $alt_svc always;   # empty value: header is not sent
    add_header Set-Cookie "h3c=$h3_cohort; Path=/; Max-Age=86400" always;   # cohort for RUM
}

To roll back, stop advertising and, for clients that already cached the advertisement, send Alt-Svc: clear, which tells them to forget alternatives for the origin. Keep the QUIC listener running for at least one max-age after you stop advertising, so cached clients get a working path rather than a timeout, unless the listener itself is the problem.

Measuring: adoption, failure and speed

Start with the server's view. nginx exposes the negotiated protocol in the $http3 variable, which is h3 for HTTP/3 connections and empty otherwise; log it next to $request_time and the cohort.

# nginx.conf
log_format h3 '$remote_addr "$request" $status $request_time '
              'proto=$server_protocol h3=$http3 cohort=$h3_cohort';
access_log /var/log/nginx/access_h3.log h3;

# share.py: per-cohort HTTP/3 share and latency percentiles
import re, statistics, sys
rx = re.compile(r'" \d+ (?P<t>[\d.]+) proto=\S+ h3=(?P<h3>\S*) cohort=(?P<c>\S*)')
rows = [m.groupdict() for m in map(rx.search, open(sys.argv[1])) if m]
for cohort in ("on", "off"):
    sel = [r for r in rows if r["c"] == cohort]
    h3 = sum(r["h3"] == "h3" for r in sel)
    t = sorted(float(r["t"]) for r in sel)
    q = statistics.quantiles(t, n=100) if len(t) >= 100 else None
    print(cohort, len(sel), f"h3={h3/len(sel):.1%}" if sel else "-",
          f"p50={q[49]:.3f}s p95={q[94]:.3f}s" if q else "")

Server-side $request_time excludes the handshake and much of the client's network experience, so it is a weak speed signal. The client's view comes from real-user monitoring: PerformanceNavigationTiming.nextHopProtocol and the same field on resource timing entries report h3 when HTTP/3 was used.

const nav = performance.getEntriesByType("navigation")[0];
navigator.sendBeacon("/rum", JSON.stringify({
  proto: nav.nextHopProtocol,                    // "h3", "h2", "http/1.1"
  ttfb: nav.responseStart - nav.startTime,
  lcp: window.__lcp,                             // from a PerformanceObserver
  cohort: document.cookie.match(/h3c=(\w+)/)?.[1] || "unknown",
}));

Now the trap. Comparing users who got HTTP/3 with users who got HTTP/2 is biased: HTTP/3 users are on newer browsers and on networks that pass UDP, which are often better networks. They would look faster even if the protocol did nothing. The unbiased comparison is the advertised cohort against the holdout, regardless of which protocol each request actually used. This is an intent-to-treat comparison: it includes the clients that tried and fell back, which is exactly the population your decision affects.

Worked example (illustrative numbers). At 5 percent advertisement, logs show 61 percent of the cohort's requests on h3 after an hour. Naively, h3 requests have 30 percent lower p75 time to first byte than h2 requests. The cohort comparison shows 6 percent lower p75 TTFB for the advertised cohort overall, with the gain concentrated in mobile networks with high loss, and a small regression in one country where h3 share is under 5 percent. That country's ISPs throttle UDP; clients race, lose, fall back, and pay a little. The correct decision is to roll forward broadly and watch that market, not to claim a 30 percent win.

0-RTT policy

QUIC can send application data in the first flight when resuming a session, saving a round trip, but that early data can be replayed by an attacker who captured it. The safe policy is to accept 0-RTT only for requests that are harmless to repeat. RFC 8470 defines the mechanism for a terminating proxy: forward the request with Early-Data: 1 and let the origin answer 425 Too Early for anything it will not process from early data, which makes the client retry after the handshake completes.

# nginx edge: allow early data, tell the origin when a request used it
ssl_early_data on;
proxy_set_header Early-Data $ssl_early_data;   # "1" when the request arrived in early data

# origin (any framework): refuse unsafe methods sent as early data
if request.headers.get("Early-Data") == "1" and request.method not in ("GET", "HEAD", "OPTIONS"):
    return Response(status=425)

Even GET is not automatically safe: an endpoint that changes state on GET, or that is expensive enough to be a denial-of-service target when replayed, should also refuse early data.

Abuse handling across a fleet

Before a QUIC server has validated a client's address, it may send at most three times the bytes it received from that address, which limits reflection attacks using spoofed sources. Under a flood of spoofed handshakes, a server can additionally send a Retry packet carrying a token the client must echo, proving it can receive at its claimed address, at the cost of one extra round trip. In nginx quic_retry is a static on or off setting, default off; enable it if your edge sees spoofed-source floods and accept the latency cost, or rely on upstream DDoS filtering.

Retry tokens and stateless reset tokens are derived from a server secret. If each server generates its own, a token issued by one server is rejected by its neighbour, and a stateless reset sent by a server that does not know the connection cannot be authenticated by the client. nginx's quic_host_key sets that secret from a file and, by default, generates a random key, so every server and every reload differs. Distribute one key file to all servers behind the same address and rotate it deliberately.

Deploys and draining

TCP has a listening socket and an accept queue; a new process can take over new connections while the old one finishes its own. UDP has neither. During a restart, packets for existing QUIC connections can reach a process that does not own them, which will typically answer with a stateless reset, and the client must reconnect, possibly falling back to TCP. Measure resets, handshakes per second and the h3 share around each deploy; a dip at every release means your restart path is cutting connections. Mitigations depend on your stack: routing to the owning process within a host, draining hosts out of the load balancer before restarting, and short idle timeouts so drains finish quickly.

Dashboards and the kill switch

SignalSourceAlert when
HTTP/3 share in the advertised cohortaccess logs ($http3)drops sharply after a deploy or in one region
Cohort vs holdout p75 TTFB and LCPRUMadvertised cohort becomes slower
UDP receive buffer errors, packet dropshost countersrising under normal load
Stateless resets, failed handshakesserver or qlog metricsspikes at deploys
Edge DDoS drops on UDP 443edge or CDNlegitimate traffic being rate-limited

The kill switch is the advertisement, not the listener: set the cohort to 0 percent and send Alt-Svc: clear, lower the HTTPS record by removing h3, and keep the listener up until caches expire. Rehearse it once during rollout so you know how long convergence takes with your max-age and TTL.

Failure modes and trade-offs

  • Partial UDP path: one region's firewall or load balancer drops UDP; that region's h3 share collapses while others look fine. Break dashboards down by region.
  • Throttled UDP mid-connection: some networks allow handshakes then rate-limit, so connections stall instead of failing fast. Look for long tails only in the advertised cohort.
  • Small datagram paths: QUIC requires paths to carry 1,200-byte UDP payloads; tunnels with smaller limits break handshakes.
  • CPU: user-space QUIC usually costs more CPU per byte than kernel TCP with TLS. Capacity-plan the edge before going to 100 percent.
  • Visibility: encrypted transport headers defeat middlebox tooling; endpoint logs and qlog become your packet capture.
  • Trade-off: long max-age and an HTTPS record maximize adoption and first-connection benefit, but slow rollback. Short values cost a little adoption and buy control.

What to do next

  1. Add $http3 and a cohort field to access logs, and protocol plus cohort to RUM beacons, before advertising.
  2. Verify UDP 443 from outside every region and through the load balancer.
  3. Share one quic_host_key file across servers behind the same address.
  4. Advertise with Alt-Svc to 5 percent with ma=600 and keep a holdout.
  5. Compare cohort against holdout, not h3 against h2; break down by region and network type.
  6. Decide a 0-RTT policy and enforce it at the origin with 425.
  7. Rehearse the kill switch, then raise the percentage, the max-age and finally add the HTTPS record.
Key takeaway: HTTP/3 in production is a controlled experiment, not a switch. Prepare the UDP path, log the protocol, and advertise to a cohort with a short max-age so rollback converges in minutes. Judge success by comparing the advertised cohort with a holdout, because protocol cohorts are biased. Restrict 0-RTT to replay-safe requests, share token keys across the fleet, watch deploys for resets, and keep the advertisement as your kill switch.