Most LLM services authenticate to their dependencies with long-lived secrets: a model provider API key in an environment variable, a vector database password in a config map, a cloud access key baked into a CI variable. Each secret is a bearer credential. Whoever reads it can use it from anywhere until someone notices and rotates it, and LLM systems add new ways to read it: agents that execute code, tools that echo their environment, traces and prompts that capture headers.
Workload identity replaces those secrets with something different in kind. The platform that runs your code vouches for what the code is, issues it a short-lived, audience-bound credential, and your dependencies trust the platform rather than a shared string. This article explains the model from first principles, shows how it works on Kubernetes and the major clouds, applies it to a typical LLM stack of gateway, retriever, tools and model provider, and covers the LLM-specific traps, chief among them agents that can steal the very identity they run under.
Identity, credential and attestation
Separate three ideas that secrets conflate. An identity is a name for a workload, such as spiffe://prod.example.com/ns/rag/sa/retriever or the Kubernetes subject system:serviceaccount:rag:retriever. A credential is proof of that identity presented to someone else. Attestation is how the issuer decides the workload deserves the identity in the first place: the kubelet knows which pod is asking, a cloud knows which VM, a SPIRE agent inspects the process and its pod.
With a static API key, all three collapse into one string that anyone can copy. With workload identity, the credential is minted by an issuer that attested the caller, it expires in minutes to hours, and it names an audience, the one service it is meant for. Stealing it is still possible, but the stolen token is useful only briefly and only against one recipient.
| Property | Static API key | Workload identity credential |
|---|---|---|
| Lifetime | until rotated, often months | minutes to hours, refreshed automatically |
| Who issues it | a human or a pipeline, once | the platform, after attesting the workload |
| Where it works | any network, any caller | one audience; often bound to a TLS key |
| Rotation | manual, coordinated, often skipped | continuous by design |
| Audit answer to 'who called?' | whoever holds the key | a specific workload name |
How the flow works on real platforms
The flow is the same everywhere, with different names. The platform injects a signed token describing the workload. The workload presents it either directly to a peer that trusts the issuer, or to a token service that exchanges it for credentials the target understands. The diagram shows this for a retrieval-augmented chatbot.
On Kubernetes the primitive is the projected service account token: a JWT the kubelet writes into the pod, with a configurable audience and expiry, signed by the cluster's issuer and refreshed automatically before it expires. Everything else builds on it:
- AWS IRSA projects a token whose audience is
sts.amazonaws.comand setsAWS_ROLE_ARNandAWS_WEB_IDENTITY_TOKEN_FILE. The AWS SDK callsAssumeRoleWithWebIdentity; STS checks the signature against the cluster's OIDC issuer and matches the token's subject and audience against the role's trust policy. - EKS Pod Identity uses a node agent instead: the SDK asks the agent, the agent exchanges a token with audience
pods.eks.amazonaws.com, and the same role can serve several clusters without adding each cluster's OIDC provider to the trust policy. - GKE Workload Identity Federation exchanges the pod's token through Google's Security Token Service for a short-lived access token, and IAM policy grants Google Cloud permissions to the Kubernetes identity.
- SPIFFE and SPIRE give each workload a SPIFFE ID and issue X.509-SVIDs for mutual TLS or JWT-SVIDs for bearer use through a local Workload API socket, independent of any cloud. This is the usual choice for service-to-service authentication inside the stack.
Choosing between them is mostly a question of where the caller runs and what the recipient understands:
| Caller runs on | Recipient | Use | Watch out for |
|---|---|---|---|
| EKS | AWS APIs, including Bedrock | EKS Pod Identity or IRSA | trust policy subject pinning; one role per component |
| GKE | Google Cloud APIs, including Vertex AI | Workload Identity Federation for GKE | granting roles to a whole namespace by accident |
| Any Kubernetes or VMs | your own internal services | SPIFFE SVIDs via SPIRE or a service mesh | SVID rotation reaching every sidecar; trust-domain naming |
| CI pipeline | cloud APIs for evaluation or deploy jobs | the CI system's OIDC token federated to a cloud role | subject conditions that match any branch or any repository |
| Anywhere | a vendor that accepts only API keys | a broker holding the key, authenticated by workload identity | the broker becoming a shared, unaudited proxy |
The CI row is easy to forget. Evaluation pipelines that call model APIs with stored keys are a common leak source, and most CI platforms can now issue OIDC tokens that a cloud role trusts directly. Hardening the rest of the deployment is covered in LLM deployment hardening.
Applying it to an LLM stack
Apply the model to each hop of a typical LLM application:
- Service to model provider. Cloud-hosted model APIs accept cloud credentials: Vertex AI accepts Google OAuth access tokens governed by IAM roles, Bedrock accepts SigV4-signed requests authorized by IAM actions such as
bedrock:InvokeModel, and Azure OpenAI accepts Microsoft Entra ID tokens. Each fits workload identity directly, with no API key in the pod. - Service to service. Gateway, retriever, reranker and tool servers authenticate each other with mTLS using SVIDs or a service mesh, and authorize by peer identity: only
rag/retrievermay query the vector index. The mechanics of peer authorization and a permissive-to-strict rollout are covered in mTLS for internal services; this page focuses on what changes when one of the peers is an LLM component. - Service to API-key-only vendors. Some providers accept only static keys. Do not spread those keys around. Put the key in exactly one broker, usually the LLM gateway, loaded from a secrets manager by the gateway's own workload identity; every other component calls the gateway and authenticates with its own identity. The key then has one reader, and every model call is attributed to a workload.
- Agent acting for a user. Workload identity proves which service is calling, not on whose behalf. When an agent calls a tool for a user, carry both: the workload identity on the connection and a delegated user token, exchanged with OAuth token exchange (RFC 8693) so the tool sees an actor and a subject. The authorization rule for combining them is covered in the confused deputy in depth.
Wiring and verification code
Here is the IRSA wiring for a gateway that calls Bedrock. The trust policy pins the exact service account and audience; nothing else in the account or cluster can assume the role.
# Kubernetes: the gateway's service account, annotated with its role
apiVersion: v1
kind: ServiceAccount
metadata:
name: gateway
namespace: rag
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::111122223333:role/rag-gateway
# IAM trust policy on rag-gateway (condition block)
"Condition": {
"StringEquals": {
"oidc.eks.REGION.amazonaws.com/id/CLUSTER_ID:sub": "system:serviceaccount:rag:gateway",
"oidc.eks.REGION.amazonaws.com/id/CLUSTER_ID:aud": "sts.amazonaws.com"
}
}The application code contains no credentials at all. The default credential chain finds the web identity token file, exchanges it and refreshes as needed:
import boto3, json
bedrock = boto3.client("bedrock-runtime") # creds from the projected token
resp = bedrock.invoke_model(modelId=MODEL_ID, body=json.dumps(payload))Internal services that accept JWTs from peers must verify them properly. The checks are short, and skipping any one of them is a real vulnerability:
import jwt # PyJWT
def verify_peer(token, jwks_client, expected_aud):
key = jwks_client.get_signing_key_from_jwt(token).key
claims = jwt.decode(
token, key,
algorithms=["RS256", "ES256"], # never accept "none"
audience=expected_aud, # this service, not a wildcard
issuer="https://oidc.example.internal", # the one issuer you trust
options={"require": ["exp", "iat", "sub", "aud"]},
leeway=30, # small clock-skew allowance
)
if claims["sub"] not in ALLOWED_CALLERS: # authorize the workload name
raise PermissionError(claims["sub"])
return claims
LLM-specific traps
LLM systems break workload identity in ways ordinary microservices rarely do, because part of the system executes instructions that arrive as data.
- Code-executing agents inherit the pod's identity. If the agent's Python sandbox runs inside the gateway pod, generated code can read the projected token file or query the cloud metadata endpoint and walk away with the gateway's permissions. Run sandboxes in separate pods with
automountServiceAccountToken: false, no IRSA annotation, and network policy that blocks the metadata address and internal services. - Tools that echo their environment. A diagnostic tool that returns environment variables or file contents can hand a token file path, or the token itself, to the model and from there into logs or the user's screen. Tools should never read credential paths, and outputs should be scanned for token-shaped strings.
- Tokens in traces and prompts. LLM observability stacks capture full request payloads. Redact
Authorizationheaders and JWT-shaped strings before they reach trace storage; a prompt log is a long-lived copy of a short-lived token. - One identity for everything. Sharing a service account between the gateway, the retriever and the agent runtime means the least-trusted component holds the most-trusted permissions. Give each component, and ideally each agent type, its own identity.
- Tenant isolation by prompt. In multi-tenant RAG, tenant boundaries enforced only by the retriever's query filter fail open on a bug. Where the data store supports it, authorize per-tenant indexes against identity, as discussed in tenant isolation for LLM systems.
Worked example: migrating a chatbot off static keys
A team runs a support chatbot: a gateway pod holding an OpenAI-style API key in an environment variable, a retriever with a vector database password, and an agent that can run Python for data questions. The migration takes four steps, each deployable on its own.
- Split the code sandbox into its own deployment with no service account token, no cloud role and egress only to an allowlisted package mirror. This closes the worst path first: generated code can no longer read anything worth stealing. See egress control.
- Issue SVIDs through SPIRE or the existing service mesh, and switch the vector database to mTLS client authentication that accepts only the retriever's ID. Remove the password.
- Move the model call to a cloud-hosted endpoint reachable through IRSA or Workload Identity Federation, or, if the vendor accepts only keys, move the key into a secrets manager readable solely by the gateway's role, with rotation every 30 days.
- Turn on audit logs that record the workload identity of every model, database and tool call, and alert on any identity calling a resource it has never called before.
After the migration the only static secret is the vendor key, readable by one identity and logged on every use. A prompt injection that achieves code execution now lands in a pod with nothing to steal and nowhere to send it. For the audit layout, see audit logging for LLM systems.
Operating it
Operational rules that keep the system healthy:
- Set token audiences per recipient and never reuse one audience across services; a token for the vector database must not be accepted by the tool server.
- Make clients re-read the token file or use the SDK's credential provider. A client built once from the file contents at startup fails the first time the token expires, typically hours after deploy, which looks like a random outage.
- Keep clocks synchronized and allow seconds, not minutes, of leeway.
- Pin trust policies to exact subjects. A wildcard subject such as
system:serviceaccount:*:*lets any pod in the cluster assume the role. - Monitor issuance failures. If the issuer or STS is unreachable, workloads keep cached credentials until expiry and then fail together; alert on refresh errors well before that.
- Inventory remaining static secrets with a scanner and track their count down to the brokered minimum.
What to do next
- List every credential your LLM services hold today, its reader, its lifetime and its last rotation.
- Move any code-execution sandbox into a separate pod with no service account token, no cloud role and blocked metadata access.
- Give the gateway, retriever, tool servers and agent runtime separate service accounts and separate cloud roles.
- Switch cloud-hosted model calls to IRSA, EKS Pod Identity, Workload Identity Federation or Entra workload identities, and delete the keys.
- Adopt mTLS with SPIFFE IDs or your mesh for internal hops, and authorize by peer identity.
- Fence each remaining vendor API key inside one broker, loaded from a secrets manager, with rotation and per-call attribution.
- Add trace redaction for tokens and an alert for any identity calling a new resource.