Zero trust is usually explained as a slogan, never trust, always verify, followed by a list of pillars. That is a fine starting point, and the zero trust architecture overview on this site covers the pillars. This article is about the part the slogan skips: where, concretely, the verification happens in a cloud estate, what code and policy you have to write, and which gaps attackers use when a team believes it has adopted zero trust but has only moved the VPN.
The core idea fits in one sentence. Access is granted per request, to a specific resource, based on an authenticated identity and the context of that request, and never because the request arrived from a trusted network. In a cloud, that sentence has to be realised at four different kinds of boundary: people reaching applications, workloads reaching cloud APIs, services reaching each other, and anything reaching data. Each has different tools, and a design is only as strong as the weakest of the four.
The model: decision points and enforcement points
NIST SP 800-207 describes a zero trust architecture in terms of a policy decision point, which evaluates a request against policy and signals such as identity, device health and risk, and a policy enforcement point, which sits in the request path and allows, denies or terminates the session as told. The standard's useful insight is structural: the enforcement point must be in front of the resource, and everything behind it must be reachable only through it. Draw your system as a set of resources, and for each one ask two questions. What enforces access in front of it? And is there any path to it that avoids that enforcer?
Cloud platforms give you enforcement points at every layer, though they are rarely named that way. Cloud IAM is the enforcement point for every control-plane and data-plane API call. An identity-aware proxy is the enforcement point for web applications. A service mesh sidecar or an mTLS-terminating library is the enforcement point between services. Resource policies on storage buckets, keys and databases are the enforcement point for data. The work of zero trust in the cloud is making each of these decide on identity rather than on network origin, and removing the bypass paths.
People to applications: the identity-aware proxy
The pattern Google published as BeyondCorp puts every internal web application behind a proxy that authenticates the user with the corporate identity provider, evaluates device and context signals, and only then forwards the request. Managed versions exist on the major clouds: Google Cloud Identity-Aware Proxy in front of load balancers, App Engine and Cloud Run, and AWS Verified Access in front of load balancers and network interfaces, with policies evaluated per request against identity and device trust providers. Access policies say which groups may reach which application under which conditions, for example managed device and a recent strong authentication.
The proxy must tell the application who the user is, and it does so with a signed assertion. IAP adds a header named x-goog-iap-jwt-assertion containing an ES256-signed JWT whose issuer is https://cloud.google.com/iap and whose audience identifies the backend service, App Engine app or Cloud Run service. Verified Access passes user claims in x-amzn-ava-user-context as an ES384-signed JWT. In both cases the application should verify the signature, expiry, issuer and audience before trusting any identity field. Reading an unsigned email header, which some older proxies set, is not verification: anything that can reach the application can forge it.
# Flask backend behind Google Cloud IAP. The proxy adds x-goog-iap-jwt-assertion (ES256).
from flask import Flask, request, abort
from google.auth.transport import requests as grequests
from google.oauth2 import id_token
AUDIENCE = "/projects/123456789012/global/backendServices/987654321" # yours, from the console
IAP_CERTS = "https://www.gstatic.com/iap/verify/public_key"
app = Flask(__name__)
_session = grequests.Request()
def iap_identity():
token = request.headers.get("x-goog-iap-jwt-assertion")
if not token:
abort(401) # traffic did not come through the proxy
try:
claims = id_token.verify_token(token, _session, audience=AUDIENCE,
certs_url=IAP_CERTS)
except Exception: # bad signature, expired, wrong audience
abort(401)
if claims.get("iss") != "https://cloud.google.com/iap":
abort(401)
return claims["sub"], claims.get("email")
@app.route("/admin/refunds", methods=["POST"])
def refund():
sub, email = iap_identity()
# Proxy said "may reach the app". The app still decides "may do THIS":
if not is_authorised(sub, action="refund.create", amount=request.json["amount"]):
abort(403)
return do_refund(request.json, actor=email)Note the split in that handler. The proxy answered a coarse question: may this person, on this device, reach this application at all? The application answers the fine one: may this person issue this refund of this amount? Zero trust does not move authorisation out of applications; it guarantees that the identity they authorise is real.
Closing the bypass paths
The most common zero trust failure is an application that verifies the proxy's identity but is also reachable without the proxy. A backend with a public IP, a load balancer that still listens on another hostname, a debug port opened in a firewall rule last year, or a peered VPC where every host can reach it directly: each turns the proxy into a suggestion. Remove public addresses from backends. Restrict ingress to the proxy's or load balancer's documented source ranges, or better, to private connectivity only. Treat the header verification as mandatory rather than optional, so that a request that did arrive by another route is rejected anyway. Test it: from a host inside the VPC, call the backend directly and confirm a 401.
Workloads to cloud APIs: no static keys
A long-lived access key in a CI system is the opposite of zero trust: it authenticates whoever holds it, from anywhere, for months. Workload identity federation replaces it. The CI system issues a short-lived OIDC token describing the job (repository, branch, environment), the cloud's security token service validates that token against a trust policy, and returns credentials that expire in minutes. On AWS, a GitHub Actions job assumes a role through a trust policy that pins both the audience and the subject claim; Google Cloud's workload identity federation and Azure's federated credentials implement the same exchange.
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {"Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com"},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": "repo:example-org/payments:ref:refs/heads/main"
}
}
}]
}The subject condition is the security boundary. Pin it to a repository and a protected branch or deployment environment; a wildcard such as repo:example-org/* lets any repository in the organisation, and any pull request branch within it, assume a production role. Inside the cloud, the same principle means using the platform's attached identities (instance roles, GKE workload identity, managed identities) instead of key files, and keeping keys that must exist in a managed key service such as Cloud KMS, where their use is itself an audited, policy-checked API call.
Services to services: identity on every connection
Inside a cluster, a request from the payments service to the ledger service should carry the caller's identity, not just its IP address. Mutual TLS gives each workload a short-lived certificate encoding a workload identity, commonly a SPIFFE ID of the form spiffe://trust-domain/ns/payments/sa/api, and each side verifies the other during the handshake; the mTLS guide covers the handshake and certificate rotation. Encryption alone is not the goal. The enforcement point is the authorisation policy evaluated after the handshake: the ledger accepts writes only from the payments identity, reads from the reporting identity, and nothing from anyone else, whatever network the packet crossed.
Service meshes provide this with sidecars or node proxies and a policy language; without a mesh, the same result comes from a TLS library configured to require client certificates and a small middleware that maps the verified peer identity to allowed operations. Either way, network segmentation remains useful as a second layer that limits blast radius, but it stops being the thing that grants access.
Anything to data: private endpoints and resource policies
Managed data services are public APIs by default: a storage bucket is addressable from the internet, and only IAM stands between it and the world. Zero trust for data combines two checks. The principal must be allowed, and the request must arrive over a path you control. Private endpoints put the service on a private address inside your network, and resource policies can then refuse any request that did not use that endpoint, even from a principal that is otherwise allowed. Leaked credentials used from an attacker's laptop then fail.
{
"Sid": "DenyUnlessViaOurEndpoint",
"Effect": "Deny",
"Principal": "*",
"Action": "s3:*",
"Resource": ["arn:aws:s3:::payments-ledger", "arn:aws:s3:::payments-ledger/*"],
"Condition": {"StringNotEquals": {"aws:SourceVpce": "vpce-0abc123def4567890"}}
}Google Cloud expresses the same idea with VPC Service Controls perimeters around projects and services, and Azure with private endpoints plus disabling public network access on the resource. Roll such policies out carefully: an explicit deny also blocks the console, backup tools and cross-account replication unless you carve them out, so start in a logging or dry-run mode where the platform offers one.
Continuous verification in practice
Per-request decisions only matter if the signals are fresh. Keep sessions and tokens short: proxy sessions measured in hours, federated credentials in minutes, workload certificates in hours or less. Feed device posture from your endpoint management system into proxy policy so a laptop that falls out of compliance loses access at its next request, not at its next login. Require step-up authentication for sensitive paths such as production consoles. And log every decision, allow and deny, with principal, device, resource and the policy version that decided, in one store you can query during an incident. Without the logs you cannot tell whether a policy is too loose, and you cannot safely make it tighter.
Worked example: retiring the VPN for an operations console
A company runs an internal operations console on Kubernetes. Today it is reachable by anyone on the corporate VPN, it trusts an X-User header set by an old reverse proxy, and its deploy pipeline uses a static cloud key with administrator rights.
Step one puts the console behind the cloud's identity-aware proxy, with an access policy allowing the operations group from managed devices. The application gains the verification function shown above and stops reading X-User. Step two removes the service's external address and restricts ingress to the load balancer path; a test from a VPC host now gets 401. Step three moves authorisation for dangerous actions into the application: refunds above a threshold require a second approver identity recorded in the audit log. Step four replaces the pipeline key with OIDC federation pinned to the main branch of the deploy repository, scoped to a role that can update this one deployment. Step five puts the console's database behind a private endpoint with a policy that denies access from anywhere else.
After the change, a stolen VPN credential is useless, a leaked pipeline token expires in minutes and only works for one repository and branch, and a compromised pod elsewhere in the cluster cannot call the console's backend because it has neither the proxy assertion nor the mTLS identity the console accepts. Nothing in that list depends on which network the attacker is on.
Trade-offs and failure modes
- Proxy as single point of failure. If the identity provider or proxy is down, every application is down. Keep a documented break-glass path with separate credentials, strong alerting and after-the-fact review.
- Unverified headers. Trusting identity headers without checking signatures, or checking signatures without checking audience, lets one application's token be replayed against another.
- Wildcard federation. Loose subject conditions on OIDC trust policies are the keyless version of a shared admin key.
- Coarse proxy, no app authorisation. Everyone who may reach the app can do everything in it. The proxy is the front door, not the permission system.
- Deny policies without a rollout plan. Endpoint-bound resource policies break backup, analytics and support tooling that nobody listed. Inventory callers from access logs first.
- Latency and cost. Per-request checks, extra hops and short-lived credentials add milliseconds and token-service calls. Cache verification keys, reuse sessions, and measure rather than guess.
What to do next
- List every internet-reachable and VPN-reachable application and record its enforcement point; any entry that says network is a gap.
- Put one internal web application behind an identity-aware proxy, verify the signed assertion in code, and prove from inside the VPC that direct access fails.
- Search your CI systems and repositories for long-lived cloud keys and replace each with OIDC federation pinned to repository and branch or environment.
- Turn on mTLS with an authorisation policy for one sensitive service pair and confirm an unlisted caller is rejected.
- Put your most sensitive bucket or database behind a private endpoint with an endpoint-bound resource policy, starting in dry-run mode where available.
- Send all allow and deny decisions to one searchable store and review the denies weekly.