Identity and access management answers four questions on every request: who is calling, how do we know, what are they allowed to do, and who can change that answer. Most breaches involving cloud accounts are failures of one of those four, not of cryptography: a long-lived key in a public repository, a trust policy that accepts any repository in an organisation, a role that can pass itself a more powerful role. IAM feels like configuration, but it behaves like a distributed system with an evaluation engine at its centre, and it rewards being understood as one.

This article treats IAM as architecture. It lays out the components, separates human from workload identity, walks through how a cloud policy engine evaluates a request, explains the guardrail layers that bound what any grant can do, and works through a keyless GitHub Actions deploy to AWS end to end. Authorization models themselves, roles versus attributes, are covered in RBAC vs ABAC architecture; here the question is how the whole machine fits together.

Advertisement

The components

Every IAM system, whether a cloud provider's, Kubernetes' or your own application's, has the same parts. An identity provider authenticates principals and issues assertions: SAML assertions or OIDC ID tokens for people, signed tokens for workloads. A token service exchanges those assertions for credentials that the resource APIs accept, after checking a trust policy. An enforcement point, the API front door, SDK middleware or a service-mesh sidecar, intercepts each request and asks a decision point whether to allow it. The decision point evaluates the request's principal, action, resource and context against policies held in a policy store. A directory, ideally fed from the HR system, drives the lifecycle, and an audit log records every decision.

Two separations matter. Authentication happens rarely, at the identity provider; authorization happens on every call, at the decision point. And the policy store holds two kinds of policy: grants, which allow things, and guardrails, which bound what grants may allow. Keeping those apart is what lets a central security team delegate role creation to product teams without losing control.

IAM as a system: identity is proven once, credentials are short-lived, every call is decided and loggedHuman / workloadperson, pod, CI jobIdentity providerSSO, MFA, OIDC issuerToken servicetrust policy, STSShort-lived credsminutes to hoursEnforcement pointAPI gateway, SDK, sidecarDecision pointevaluate all policiesDirectory / HRjoiner, mover, leaverPolicy storeidentity, resource, guardrailsAudit logwho, what, allowed?authenticateassertion / JWTissuesigned requestcontextSCIMlogAuthentication happens at the identity provider; authorization happens on every request at the decision point.Guardrail policies in the store bound what any grant can do, whoever wrote it.
IAM components and the path of one request. The token service and decision point are the two places a design mistake becomes a breach.

Principals: people, workloads and federation

Human principals should never hold long-lived cloud credentials. They sign in to a workforce identity provider with phishing-resistant MFA (see WebAuthn for why passkeys resist phishing), and federate into cloud accounts through SAML or OIDC, receiving session credentials that expire in hours. Access to production is granted through roles assumed for a task, ideally just in time, which privileged access management covers in depth.

Workload principals are now the majority in most estates and the most neglected. A pod, a function, a CI job or a VM each needs an identity. The anti-pattern is a static access key stored in a secret; the pattern is workload identity federation: the platform the workload runs on (Kubernetes, the cloud's instance metadata service, GitHub Actions) issues a signed, short-lived token describing the workload, and the cloud's token service exchanges it for credentials after checking the token's issuer, audience and subject against a trust policy. Nothing long-lived exists to leak.

The trust policy is therefore the most security-critical document in the system. It decides which external identities may become which internal role, and it is evaluated only at credential issuance, so a mistake there is invisible in day-to-day permission reviews.

Advertisement

How a policy engine decides

Cloud policy engines are default-deny and deny-overrides. AWS's documented evaluation logic is a good concrete model because it has every layer. For a single call the engine gathers every applicable policy: identity-based policies on the principal, the resource-based policy on the target, and the guardrails in force, which are service control policies (SCPs) and resource control policies (RCPs) from AWS Organizations, a permissions boundary on the principal, and a session policy passed when the role was assumed. Then it applies the rules below.

def decide(request, policies):
    """Simplified AWS-style evaluation for one API call. Default is deny."""
    applicable = [p for p in policies if p.matches(request)]    # action, resource, conditions

    # 1. An explicit Deny anywhere wins, whatever else allows the call.
    if any(p.effect == "Deny" for p in applicable):
        return "DENY (explicit)"

    # 2. Guardrails must each allow; they grant nothing on their own.
    for layer in ("scp", "rcp", "permissions_boundary", "session_policy"):
        if request.has_layer(layer) and not any(
                p.layer == layer and p.effect == "Allow" for p in applicable):
            return f"DENY (not allowed by {layer})"

    identity_ok = any(p.layer == "identity" and p.effect == "Allow" for p in applicable)
    resource_ok = any(p.layer == "resource" and p.effect == "Allow" for p in applicable)

    # 3. Cross-account: both sides must allow. Same account: either may suffice
    #    (with documented exceptions, e.g. KMS key policies and IAM role trust policies).
    if request.cross_account:
        return "ALLOW" if identity_ok and resource_ok else "DENY (implicit)"
    return "ALLOW" if identity_ok or resource_ok else "DENY (implicit)"

Three properties follow. An explicit deny in any layer ends evaluation, which is why guardrails are usually written as denies with conditions. Guardrail layers never grant; a permissions boundary that allows s3:* gives nothing to a role whose identity policy allows only s3:GetObject. And cross-account access needs consent from both sides: the calling account's identity policy and the target's resource policy. The code above is a simplification, the real engine has service-specific exceptions, so treat it as the mental model and use the provider's simulator for actual answers.

Guardrail layers

Guardrails exist so that a central team can bound a large, delegated estate. SCPs set the maximum permissions for every principal in an account or organisational unit, including the account's own administrators; typical uses are denying regions you do not operate in, denying disabling of CloudTrail, and denying any attempt to leave the organisation. RCPs do the same from the resource side for services that support them, for example denying access to your data by principals outside your organisation. A permissions boundary caps a single principal and is the tool for safe delegation: product teams may create roles only if those roles carry the platform-defined boundary, so a team cannot mint a role more powerful than the boundary allows.

Delegation has one classic hole: iam:PassRole. A principal that can pass an arbitrary role to a compute service can run code as that role, so an innocuous-looking 'deploy Lambda functions' permission becomes full administrator if the passable roles are not constrained. Always scope iam:PassRole to specific role ARNs and, where supported, to the service that may receive them.

Worked example: a keyless deploy from GitHub Actions

Team checkout deploys static assets to S3. Previously a CI secret held an access key with s3:* on all buckets; that key was two years old and had been copied into three forks. The replacement uses OIDC federation, so there is no key at all. Step one registers GitHub's OIDC issuer as an identity provider in the account. Step two creates the role checkout-deploy with this trust policy:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": {"Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com"},
    "Action": "sts:AssumeRoleWithWebIdentity",
    "Condition": {
      "StringEquals": {
        "token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
        "token.actions.githubusercontent.com:sub": "repo:acme/checkout:ref:refs/heads/main"
      }
    }
  }]
}

The sub condition is the whole security of the design: only workflows running on the main branch of acme/checkout can assume the role. A condition of repo:acme/* would let any repository in the organisation, including a contractor's sandbox, deploy checkout. Using GitHub environments and a subject bound to the environment adds a manual approval gate. Step three attaches a permissions policy that allows only s3:PutObject, s3:DeleteObject and s3:ListBucket on the one bucket. Step four updates the workflow:

# .github/workflows/deploy.yml (excerpt)
permissions:
  id-token: write      # lets the job request an OIDC token
  contents: read
jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::111122223333:role/checkout-deploy
          aws-region: eu-west-1
      - run: aws s3 sync ./dist s3://acme-checkout-assets --delete

At run time the job requests an OIDC token from GitHub, whose claims name the repository, branch and workflow. The action calls AssumeRoleWithWebIdentity; STS verifies the token's signature against GitHub's published keys, checks audience and subject against the trust policy, and returns credentials that expire, by default, after an hour. Every S3 call is evaluated against the permissions policy, any SCPs and the bucket policy, and logged in CloudTrail with the role session name. Finally, prove the negative with the policy simulator:

# Would the deploy role be allowed to delete objects in another team's bucket?
aws iam simulate-principal-policy \
  --policy-source-arn arn:aws:iam::111122223333:role/checkout-deploy \
  --action-names s3:DeleteObject \
  --resource-arns arn:aws:s3:::acme-ledger-exports/2026/report.csv
# Expected EvalDecision: implicitDeny. Anything else is a finding.

Delete the old access key only after CloudTrail shows no use of it for a full release cycle, and keep that check automated, because forgotten keys are how migrations end with two ways in instead of one.

Lifecycle and verification

Permissions accumulate unless something removes them. Drive joiner, mover and leaver events from the HR system into the identity provider and on to downstream apps with SCIM, so that a leaver loses every account in minutes and a mover loses their old team's access rather than keeping it. Grant humans access through groups mapped to roles, never directly, so reviews examine a handful of group memberships instead of thousands of individual grants.

Verify continuously rather than annually. Use access-analysis tooling to find resources shared outside the organisation and permissions unused for 90 days, and remove them. Run policy checks in CI on infrastructure code so that a wildcard action or a broad trust policy fails review before it merges. Keep a break-glass role, protected by hardware MFA and alerting on every use, for the day the identity provider is down. Feed authorization failures and privilege changes into your SIEM; a spike in access-denied errors from one principal is often the first sign of stolen credentials being explored.

Failure modes

  • Over-broad trust. OIDC subjects with wildcards, or cross-account roles that trust an entire account; the grant looks narrow while the set of callers is huge.
  • Confused deputy. A third-party service assumes a role in your account on behalf of any of its customers. Require an external ID and conditions such as aws:SourceAccount or aws:SourceArn.
  • PassRole escalation. Unscoped iam:PassRole plus any compute service equals administrator.
  • Long-lived keys. Static keys leak through logs, laptops and forks. Replace them with federation and alert on any key older than 90 days.
  • Revocation lag. Credentials already issued stay valid until expiry, and IAM changes are eventually consistent. Keep session durations short and know your provider's session-revocation mechanism.
  • Guardrail blind spots. SCPs do not apply to the organisation's management account; keep workloads out of it.

Trade-offs

Fine-grained, least-privilege policies are safer but costly to write and maintain; teams that over-tighten get blocked deploys and start requesting wildcards. The practical balance is broad guardrails centrally, narrow grants generated from observed usage, and fast, audited just-in-time elevation for the rare cases in between. Short sessions reduce the value of stolen credentials but increase token-service traffic and re-authentication friction. Central policy engines give consistency but become a dependency on every request, which is why cloud providers evaluate locally in each service and why your own application authorization should cache decisions carefully rather than call out per request.

What to do next

  1. Inventory every long-lived credential, human and workload, and set a date to replace each with federation.
  2. Read every role trust policy and remove wildcard subjects and account-wide trusts.
  3. Scope every iam:PassRole grant to specific role ARNs.
  4. Put SCPs in place for region restriction, audit-log protection and leaving the organisation.
  5. Require permissions boundaries on any role created by delegated teams.
  6. Automate unused-permission removal, run the policy simulator in CI for critical denies, and alert on break-glass use.
Key takeaway: IAM is a system with a clear shape: an identity provider proves who is calling, a token service turns that proof into short-lived credentials under a trust policy, and a decision point evaluates every request against grants and guardrails with default-deny and deny-overrides. Put most of your care into trust policies and PassRole, because they decide who can become what. Replace static keys with federation, bound delegated teams with SCPs and permissions boundaries, drive lifecycle from HR through SCIM, and verify continuously with simulators, access analysis and audit logs.