Most explanations of website hosting stop at "put the files in a bucket and add a CDN". That gets a page online, but it does not explain why a deploy can leave users on a blank screen, why the apex domain will not accept a CNAME, or why a missing file comes back as 403. Each of these problems comes from how the layers interact, not from any single product.
This article splits the stack into five planes, shows how to choose between the three common architectures, follows one request from browser to origin, and then concentrates on what decides whether users ever see a broken page: cache policy and the deploy pipeline. The examples use AWS names (S3, CloudFront, Route 53, ACM), but every idea maps onto Google Cloud, Azure or a commercial CDN.
The five planes of a hosting stack
A website that serves real traffic is five cooperating systems. Treating them separately makes both design and debugging easier, because every symptom belongs to exactly one plane.
- Naming. The registrar that owns the domain, the authoritative DNS provider it delegates to, and the records that point the name at the edge. This plane changes rarely, but mistakes here take hours to undo because resolvers cache them.
- Edge. The CDN points of presence that end TLS, apply the web application firewall, answer from cache, and route cache misses to an origin by path. Almost every request should end here.
- Origin. Where content really lives: an object store for files, an API service for dynamic data, and optionally a server-side rendering (SSR) service. Origins should be reachable only from the edge.
- Build and deploy. The pipeline that turns source into files with content-hashed names, uploads them in a safe order, and tells the edge what changed.
- Observation. Edge logs, origin metrics, synthetic checks from outside, and certificate and domain expiry alarms. Without this plane you learn about outages from users.
Choosing the architecture: static, static plus API, or server-rendered
The first decision is how much of each page is computed per request. Push as much as possible toward the static end, because every step toward dynamic rendering adds a service to run, a cache key to get right and a way to fail.
| Architecture | What runs per request | Good for | Main cost |
|---|---|---|---|
| Static | Nothing. The edge serves files from cache or object storage | Documentation, blogs, marketing, most single-page apps | Every content change needs a build and deploy |
| Static plus API | Only the API calls a page makes after it loads | Dashboards and apps whose shell is the same for every user | Two deploy units and API versioning between them |
| Server-rendered | HTML is generated per request, maybe cached briefly | Personalised pages, very large catalogues, SEO on data that changes fast | Compute at the origin, cache keys based on cookies, cold starts |
A useful rule: if two users asking for the same URL at the same moment should get the same bytes, the page is static no matter how often it changes. A news home page rebuilt every five minutes is static; a shopping cart is not. Most real sites are hybrids, with pages under /, an API under /api/, and one CDN routing each path pattern to its origin with its own cache policy.
The request path, hop by hop
When a reader opens https://example.com/guide/, the browser first resolves the name. The recursive resolver asks the authoritative servers for example.com, which answer with the CDN's addresses. The browser then connects to the nearest edge location, completes a TLS handshake using the certificate attached to the distribution, and sends the request. HTTP/3 over QUIC removes a round trip where both sides support it.
The edge builds a cache key from the host, the path and whatever query strings, headers and cookies the cache policy includes. On a hit it answers straight away. On a miss it forwards the request to the origin selected by the first path pattern that matches, stores the response as its headers allow, and returns it. The origin sees only misses, so its load depends on the hit ratio, not on total traffic. A static site with a 98 percent hit ratio sends the origin one request in fifty.
Two details trip people up. First, the edge uses the origin's Cache-Control headers, within the minimum and maximum TTLs of its cache policy, to decide how long to keep an object, so the headers you set at upload time are part of your architecture. Second, edge addresses change and differ by location, so always point DNS at the CDN by name.
Naming: the apex problem and TTL strategy
DNS does not allow a CNAME at the zone apex (example.com itself), because the apex must also hold SOA and NS records and a CNAME cannot share a name with other records. Yet CDNs want you to point at their hostname. Providers solve this in one of three ways: an alias record that the DNS provider resolves internally (Route 53 alias, which can target a CloudFront distribution), CNAME flattening at resolve time, or static anycast addresses that you put in a plain A record. Whichever you use, redirect one of www and the apex to the other so there is one canonical host.
Pick TTLs by how quickly you might need to change a record. Records pointing at a CDN rarely change, so an hour or more is fine. Before a migration, lower the TTL, wait one full old TTL, make the change, and raise it again. The parent zone's NS records often have TTLs of a day or more, so a change of DNS provider must serve identical records from both providers until the old delegation has expired.
TLS and origin lockdown
Certificates are issued for names, so the certificate on the distribution must cover every hostname you serve, usually the apex and www. Managed certificates renew automatically, but only while the DNS validation record still exists, so never delete it. CloudFront has one rule that surprises everyone: it only accepts ACM certificates from the us-east-1 region, whatever region the rest of your stack uses. Add HSTS only once every subdomain serves HTTPS, because browsers enforce it for as long as the header's max-age says.
The origin should not be reachable from the internet at all. With S3, keep the bucket private, block public access, and let the distribution read it through origin access control (OAC), which signs the edge's requests to the bucket. The legacy origin access identity still works but is not recommended for new setups. For API and SSR origins, allow only the CDN's published address ranges, or require a secret header that the edge adds and the origin checks. Otherwise attackers can bypass the WAF and your caching by calling the origin directly.
A private bucket has one debugging surprise. When the requester may not list the bucket, S3 answers a missing key with 403 rather than 404, so it does not reveal which keys exist. A spike in 403s after a deploy usually means missing files, not broken permissions. Map these to a real 404 page. Do not map every 403 to /index.html with status 200 unless a single-page app needs it, because that hides real errors from crawlers and monitoring.
Cache policy is deploy design
The key idea of a safe static deploy is to split files into two classes. Fingerprinted assets such as app.3f9a1c.js contain a hash of their content in the name. If the content changes, the name changes, so an asset URL always means the same bytes and can be cached for a year with immutable. Entry documents, meaning HTML, robots.txt, the sitemap and any manifest, keep stable names, and they are the only files that decide which version a user sees.
For entry documents, separate the browser TTL from the edge TTL. Send Cache-Control: public, max-age=0, s-maxage=86400. Browsers revalidate on every visit, which is cheap because the edge answers with a 304. The edge keeps the page for a day. Your deploy then invalidates HTML paths so the new version appears straight away. The result is that invalidation acts on a few hundred small documents, while the large assets never need purging.
Order matters because a deploy is not atomic. Upload new assets first, then the HTML that references them, then invalidate. Reverse it and an edge can fetch new HTML pointing at an asset that does not exist yet. Deleting old assets is the other common mistake: a browser holding the previous HTML will still request that bundle's lazily loaded chunks, so keep old assets for several deploys.
Worked example: a 6,000-page static site with a pipeline
Take a documentation site with 6,000 HTML pages and about 40 MB of fingerprinted assets, hosted on S3 and CloudFront and deployed from GitHub Actions on every push to main. The script below is the whole release. Each step does exactly one thing, and the order is the safety argument from the previous section.
#!/usr/bin/env bash
set -euo pipefail
BUCKET="s3://example-docs-site"
DIST_ID="E2EXAMPLE123"
npm ci && npm run build # writes dist/, assets named by content hash
# 1. Assets first. Never --delete here: old HTML still references old chunks.
aws s3 sync dist/assets "$BUCKET/assets" --size-only \
--cache-control "public, max-age=31536000, immutable"
# 2. Entry documents second: browsers revalidate, the edge keeps them a day.
aws s3 sync dist "$BUCKET" --exclude "assets/*" --delete \
--cache-control "public, max-age=0, s-maxage=86400"
# 3. Invalidate. A trailing wildcard counts as one path.
aws cloudfront create-invalidation --distribution-id "$DIST_ID" --paths "/*"
# 4. Smoke test from outside, through the edge, not the bucket.
curl -fsS -o /dev/null -w "%{http_code}\n" https://docs.example.com/Some numbers make this concrete. A fresh CI build gives every file a new timestamp, so plain sync re-uploads everything. --size-only is safe on the asset step because asset names change with content, but never on HTML, where a same-length edit would be skipped. The wildcard invalidation also purges asset URLs, which is harmless because they are immutable. At 2 million page views a day with a 97 percent hit ratio, the bucket serves about 60,000 HTML fetches a day. A weekly job deletes asset objects older than 30 days that the current HTML does not reference.
Rollback is running the pipeline for the previous commit. Old assets are still in the bucket, so pages already open keep working throughout.
Adding dynamic parts without losing the cache
Put the API under the same host at /api/* as a second cache behaviour with caching disabled. One host means no CORS preflight, first-party cookies and one certificate. Forward only the headers and cookies the API needs. Forwarding all cookies on the static behaviour is the classic way to drop the hit ratio from 95 percent to single digits.
If you add SSR, cache anonymous responses briefly and bypass the cache when a session cookie is present. Keep personalisation in API calls rather than in the HTML, because one personalised fragment makes a whole page uncacheable.
Failure modes
- Blank page after deploy. The HTML references assets that are not at the edge yet, or that the deploy deleted. Upload assets first and keep old ones.
- Users stuck on the old version. HTML was uploaded with a long browser
max-age. Invalidation cannot reach browser caches, so the only cure is waiting. Keep the browser TTL at zero for entry documents. - Certificate expiry or validation loss. Someone removed the DNS validation CNAME, and renewal silently fails weeks later. Alarm on days to expiry from an external checker.
- Origin bypass. The bucket or API is public, so the WAF does nothing. Call the origin from outside to test.
- Hit ratio collapse. A header, cookie or query string entered the cache key. Track hit ratio per deploy.
- Domain expiry. Auto-renew failed. No architecture survives this, so enable registrar lock and renewal alerts.
Operations and trade-offs
Monitor from outside: a synthetic check that loads the home page and one deep page through the real hostname every minute, plus alerts on edge 5xx rate, origin 4xx rate after deploys, hit ratio, and certificate and domain expiry. Keep edge access logs for at least a month. Add security headers (CSP, X-Content-Type-Options, Referrer-Policy) through a response header policy at the edge so every origin gets them.
A static-first stack gives the lowest cost and smallest attack surface, and survives origin outages for cached content, but content changes wait for a build. A CDN adds a vendor dependency and a cache layer you must reason about in every incident. Multi-CDN removes the single vendor but doubles the configuration to keep in sync, and rarely pays off below very large traffic. For more detail on the individual layers, see CDN architecture in depth, Amazon CloudFront, Amazon Route 53, caching in system design and HTTP/3 and QUIC.
What to do next
- Draw your own stack as the five planes and write down the one product or script that owns each.
- Classify every file your build emits as a fingerprinted asset or an entry document, and set
Cache-Controlfor each class at upload time. - Reorder the deploy to upload assets, then HTML, then invalidate, and remove any
--deletefrom the asset step. - Add an asset cleanup job that keeps at least several releases, and test a rollback to the previous commit.
- Make the origin private: use OAC for buckets and an address allow-list or secret header for API and SSR origins, then prove it by calling the origin from outside.
- Check the apex record type, the certificate's names and region, and the DNS validation record, and add external alarms for certificate and domain expiry.
- Put the cache hit ratio and the post-deploy 4xx rate on the deploy dashboard, and treat a drop as a failed release.