Publishing a service on the internet has traditionally meant opening an inbound port: a public IP, a firewall rule, a load balancer, and the permanent exposure that comes with them. Cloudflare Tunnel inverts that. A small daemon, cloudflared, runs next to your service and makes outbound connections to Cloudflare; visitors reach a hostname on Cloudflare's network, and requests travel back down those connections to your origin. Nothing listens on a public address, so there is nothing for a port scanner to find.

This article explains how a tunnel works underneath, the two ways to manage one, how ingress rules route requests, how to run it highly available, and how to secure, monitor and debug it. It ends with a worked example, the failure modes that bite in production and a checklist. It assumes you know what DNS and a reverse proxy are, and nothing else about Cloudflare.

Advertisement

How a tunnel works

A tunnel is a named object in your Cloudflare account with a UUID. When cloudflared starts with the tunnel's credentials, it opens four long-lived connections to Cloudflare's network, spread across at least two distinct data centres, so the loss of a single connection, server or data centre does not take the tunnel down. The connections go out on port 7844. The transport is chosen by --protocol: auto (the default) uses QUIC over UDP and falls back to HTTP/2 over TCP if UDP is blocked; you can force quic or http2.

To send traffic to it, a DNS record for your hostname is a proxied CNAME pointing at <UUID>.cfargotunnel.com. When a request for that hostname reaches the edge, it passes through whatever you have configured there, such as the WAF, rate limiting or Cloudflare Access, then the edge picks one of the tunnel's live connections and multiplexes the request over it. cloudflared looks at the request, matches it against its ingress rules, and proxies it to a local service.

Visitorapp.example.comCloudflare edgeDNS, WAF, AccessTunnel IDuuid.cfargotunnel.comcloudflared replica A4 conns, 2+ data centrescloudflared replica Banother host or zoneFirewallegress 7844 onlyOrigin serviceslocalhost:3000, :8080Ingress ruleshostname, path, serviceHTTPSCNAMEoutbound QUIC / HTTP2proxied requestmatchNo inbound port is opened. The connector dials out; the edge sends requests back down those connections.Replicas add connections and availability; each request goes to the geographically closest healthy one.
Figure 1. Request path through a Cloudflare Tunnel with two replicas. The firewall only needs to allow outbound traffic to Cloudflare.

Two consequences follow. First, your origin's inbound firewall can deny everything; only egress to Cloudflare on 7844 is needed. Second, the edge is now your reverse proxy, so TLS for the public hostname terminates at Cloudflare, and the hop from cloudflared to the origin is a separate connection you configure.

Remotely and locally managed tunnels

There are two ways to own a tunnel's configuration.

  • Remotely managed. You create the tunnel in the dashboard or through the API, and the ingress rules live in Cloudflare. The connector is started with a token: cloudflared tunnel run --token <TOKEN>, or better --token-file or the TUNNEL_TOKEN environment variable so the token is not visible in the process list. This is the easiest model for containers, because the image needs no config file.
  • Locally managed. You create the tunnel with the CLI, which writes a credentials JSON file, and keep ingress rules in a config.yml next to the connector. Configuration lives in version control and changes go through review, which many teams prefer.

Tunnels are not limited to public hostnames. A tunnel can also advertise private IP ranges, for example with cloudflared tunnel route ip add 10.20.0.0/16 web-prod, so that users running Cloudflare's device client (WARP) reach internal addresses through it without a VPN concentrator. That turns the same connector into a private-network on-ramp, and the access policy, not the network path, decides who can reach what. Keep public hostname routing and private network routing in separate tunnels if different teams own them.

Treat the token or credentials file as a secret of the same class as a TLS private key: anyone who holds it can run a connector for your tunnel and receive its traffic. Store it in your secret manager, and rotate it if it leaks.

Advertisement

Ingress rules

Ingress rules map incoming requests to local services. Each rule can match a hostname and a path regular expression; cloudflared evaluates rules top to bottom and the first match wins. The last rule must match everything, usually returning 404, and the configuration is rejected without it.

tunnel: 6ff42ae2-765d-4adf-8112-31c55c1551ef
credentials-file: /etc/cloudflared/6ff42ae2-765d-4adf-8112-31c55c1551ef.json

originRequest:            # defaults for every rule
  connectTimeout: 10s     # documented default is 30s
  keepAliveConnections: 100

ingress:
  - hostname: api.example.com
    path: ^/v2/
    service: http://localhost:8081
  - hostname: api.example.com
    service: http://localhost:8080
  - hostname: grafana.example.com
    service: https://localhost:3000
    originRequest:
      originServerName: grafana.internal   # verify the origin certificate name
  - service: http_status:404               # mandatory catch-all

Order matters: the /v2/ rule must come before the bare api.example.com rule, or the broader rule captures everything. The originRequest block controls the second hop; documented defaults include a 30 second connect timeout, a 10 second TLS handshake timeout and up to 100 idle keep-alive connections. Before you deploy, let the binary check the file and tell you which rule a URL hits:

cloudflared tunnel ingress validate
cloudflared tunnel ingress rule https://api.example.com/v2/orders

Setting one up

A locally managed tunnel from nothing takes five commands:

cloudflared tunnel login                         # authorise the CLI for one zone (writes cert.pem)
cloudflared tunnel create web-prod               # creates the tunnel and its credentials JSON
cloudflared tunnel route dns web-prod api.example.com
cloudflared tunnel route dns web-prod grafana.example.com
cloudflared tunnel --config /etc/cloudflared/config.yml run web-prod

In production run it as a system service or a container, not in a terminal. On Kubernetes, a Deployment of two or more pods using a remotely managed token works well; each pod is a replica of the same tunnel. Pass the token from a Secret through TUNNEL_TOKEN, point the service at cluster DNS names, and expose the metrics port for scraping.

Replicas and high availability

A replica is another cloudflared process running the same tunnel. Each adds four connections, up to 25 replicas or 100 connections per tunnel. Run at least two, on different hosts or failure domains, so a host reboot or a connector upgrade does not cause an outage.

Replicas are not a load balancer. Cloudflare sends each request to the geographically closest healthy replica, so two replicas in the same data centre will not split traffic evenly, and you should not use replica count to scale throughput across origins. If you need weighted or health-checked distribution across several origins, put a real load balancer behind one tunnel or use Cloudflare's load-balancing product in front of several.

Two flags matter during restarts. --grace-period (default 30s) is how long the connector keeps serving in-flight requests after receiving a shutdown signal; set your orchestrator's termination grace period longer than it. --retries (default 5) bounds reconnection attempts when connections fail. Disable automatic updates in containers with --no-autoupdate and upgrade by rolling the image instead, one replica at a time.

Security

Removing the open port removes a large class of attack, but the hostname is still public. Anyone can send it requests, and they will reach your origin unless something at the edge stops them. For internal tools, put Cloudflare Access in front of the hostname so a user must authenticate before a request ever enters the tunnel, and validate the Access token at the origin as well; the zero-trust model is explained in zero trust in the cloud. For public APIs, keep the WAF and rate limits on.

Bind origin services to localhost or a private interface so the tunnel is the only path in. Avoid noTLSVerify: true except briefly while debugging; it silently accepts any certificate on the second hop. Prefer originServerName or caPool. Finally, remember that cloudflared needs only outbound access, so the host can sit in a private subnet behind a NAT gateway with no public IP at all.

Worked example: a 502 that only some users see

A team runs an API on port 8080 and Grafana on port 3000 on two VMs behind NAT. They create a locally managed tunnel with the configuration above, install cloudflared as a service on both VMs, add Access in front of Grafana, and close every inbound rule. Day one works.

A week later, api.example.com/v2/orders returns 502 while /v1 works. The connector log shows the request matched the rule for port 8081 and the dial was refused: the v2 service had been moved to a different VM and only runs there. Requests that reached the replica on the other VM failed; requests that reached the replica co-located with the service succeeded. Which replica a user hits depends on closest-replica routing from the edge location serving them, so the failures clustered by user geography and looked intermittent. The fix is to make every replica able to serve every rule, by pointing ingress at a service address reachable from both VMs, rather than assuming requests stick to one replica. cloudflared tunnel ingress rule confirmed the match in seconds; the log line identified the failed hop.

The general lesson is that a tunnel has no session affinity and no health check of your origins. Cloudflare checks that a replica's connections are alive, not that the services behind it answer. Every replica must therefore be a full copy of the routing configuration, able to reach every origin it names, and origin health belongs in your own monitoring. The team added a synthetic check per hostname from outside the network and an alert on proxy errors per replica, and the next misrouted service was caught in minutes.

Observability

Each connector exposes Prometheus metrics. Outside containers it binds 127.0.0.1 on the first free port from 20241 to 20245; in containers it binds 0.0.0.0; --metrics sets the address explicitly. Useful series include cloudflared_tunnel_ha_connections (live connections, which should be four per replica), cloudflared_tunnel_total_requests and cloudflared_tunnel_request_errors (errors proxying to the origin). Alert when HA connections drop below four times your replica count, and when request errors rise relative to requests, which almost always means the origin, not Cloudflare. Raise --loglevel from the default info to debug only while investigating.

For logs, --logfile or --log-directory send output to disk, and cloudflared tail --output=json <UUID> streams a running connector's logs as JSON you can pipe into jq. --tag attaches key=value tags to a connector, such as the host or environment, which helps tell replicas apart. When you report a problem, capture the tunnel ID, the replica, the time and the requested URL; with those four facts, the edge view and the connector log can be lined up.

Failure modes and trade-offs

The common failures and their tells:

  • No connections at start-up. Egress to 7844 is blocked. Test with a TCP connect to Cloudflare on 7844; if only UDP is blocked, auto falls back to HTTP/2 by itself.
  • 502 or 503 from the edge. The tunnel is up but the origin is unreachable from that replica: wrong port, service down, or a rule pointing at an address one replica cannot reach.
  • Requests hit the wrong service. Ingress order. Run ingress rule against the URL.
  • Dropped requests during deploys. A single replica, or an orchestrator that kills the process before the grace period ends.
  • Leaked token. Someone else can run a connector for your tunnel. Rotate it and audit connector origins in the dashboard.

The trade-offs are real. You add a dependency on Cloudflare's network and an extra proxy hop; very high-throughput or non-HTTP workloads may be better on a direct connection or private connectivity. In exchange you get no inbound exposure, free TLS and edge security, and no public IPs to manage. For failover between whole sites, combine tunnels with DNS failover and traffic steering.

What to do next

  1. Choose remotely managed (token, simplest for containers) or locally managed (config.yml in version control).
  2. Create the tunnel, route DNS for each hostname, and write ingress rules with the most specific first and a 404 catch-all last.
  3. Run ingress validate and ingress rule against your real URLs before every config change.
  4. Bind origins to localhost or a private interface and close all inbound firewall rules; allow egress on 7844.
  5. Run at least two replicas in separate failure domains, each able to reach every origin in the rules.
  6. Store the token or credentials in a secret manager and pass it with a token file or environment variable.
  7. Put Access in front of internal hostnames and keep WAF and rate limits on public ones; never leave noTLSVerify on.
  8. Scrape metrics and alert on HA connections and on request errors per replica.
  9. Disable auto-update in containers and roll upgrades one replica at a time with a termination grace longer than 30 seconds.
Key takeaway: Cloudflare Tunnel publishes services without inbound ports: cloudflared dials out on port 7844 with four connections to at least two data centres, using QUIC with HTTP/2 fallback, and the edge sends requests back down them to origins chosen by ordered ingress rules ending in a catch-all. Run two or more replicas that can each reach every origin, remembering that traffic goes to the closest replica rather than being balanced. Guard the token like a private key, put Access or the WAF in front of the hostname, verify origin TLS, and alert on HA connections and origin errors.