Amazon Cognito is the AWS service for signing users in to your applications. Most confusion about it comes from the fact that it is really two services under one name. User pools are a user directory and OpenID Connect identity provider: they store users, run sign-up and sign-in, federate with other identity providers, and issue JSON Web Tokens (JWTs). Identity pools exchange a proof of identity for temporary AWS credentials so a client can call AWS services such as S3 directly. Many applications need only the first.
This article explains how each part works, which token to use where, how to verify tokens correctly, how to customise behaviour with Lambda triggers, and which configuration choices are hard or impossible to change later. Facts were checked against the Amazon Cognito developer guide in October 2026. Prices and service quotas change and are not quoted; look them up for your Region.
The two halves and how they connect
A user pool answers the question who is this user? and expresses the answer as signed tokens your own backend can check. A user pool app client represents one application: it has an ID, optionally a secret, the OAuth flows and scopes it may use, callback URLs, and token lifetimes. Web and mobile apps that run on the user's device are public clients and must not have a secret; server-side applications can be confidential clients with one.
An identity pool answers which AWS permissions should this caller get? It accepts tokens from a user pool, from other OIDC or SAML providers, or from social providers, maps the caller to an IAM role, and returns short-lived AWS credentials. It can also issue guest credentials to unauthenticated users. The permissions are ordinary IAM policies, so everything in the IAM article applies.
Feature plans and the user pool's moving parts
User pools come in three feature plans. Lite covers basic sign-up and sign-in and the classic hosted UI. Essentials adds the newer features: the customisable managed login pages, choice-based sign-in, passwordless sign-in with passkeys, email or SMS one-time codes, email MFA and access token customisation. Plus adds threat protection, which evaluates sign-in, sign-up and password requests for signs of compromise and can log risk evaluations for export. The plan is set per pool and affects price, so check which features you actually need.
Other pieces you will configure: attributes (standard OIDC ones such as email, and custom ones prefixed custom:); a domain, which hosts managed login and the OAuth endpoints /oauth2/authorize and /oauth2/token; groups, which appear in tokens as cognito:groups; resource servers, which define custom scopes such as orders/read; and external identity providers for SAML, OIDC, and social sign-in. Federated users get a profile in the pool too, so your application sees one token format whichever provider the user came from.
Signing in: use the authorization code flow with PKCE
There are two ways to sign users in. The first is to redirect them to managed login, which runs the standard OAuth 2.0 authorization code flow. The second is to build your own screens and call the user pools API through an SDK, using flows such as Secure Remote Password (USER_SRP_AUTH), the choice-based USER_AUTH flow, or custom challenge flows. Managed login is less work and is the only way to sign in federated users; custom screens give full control over the user experience.
For browser and mobile apps, use the authorization code flow with PKCE (Proof Key for Code Exchange). The app creates a random verifier, sends its hash with the authorization request, and proves possession of the verifier when exchanging the code for tokens, so a stolen authorization code is useless without it.
import base64, hashlib, os, secrets, urllib.parse, requests
DOMAIN = "https://auth.example.com" # your user pool domain
CLIENT_ID = "1example23456789" # public app client: no secret
REDIRECT = "https://app.example.com/callback"
def start_login(session):
verifier = base64.urlsafe_b64encode(os.urandom(32)).rstrip(b"=").decode()
challenge = base64.urlsafe_b64encode(
hashlib.sha256(verifier.encode()).digest()).rstrip(b"=").decode()
state = secrets.token_urlsafe(16)
session["pkce"], session["state"] = verifier, state
q = urllib.parse.urlencode({
"response_type": "code", "client_id": CLIENT_ID, "redirect_uri": REDIRECT,
"scope": "openid email orders/read", "state": state,
"code_challenge": challenge, "code_challenge_method": "S256",
})
return f"{DOMAIN}/oauth2/authorize?{q}"
def finish_login(session, code, state):
if state != session.pop("state", None):
raise PermissionError("state mismatch")
r = requests.post(f"{DOMAIN}/oauth2/token", data={
"grant_type": "authorization_code", "client_id": CLIENT_ID,
"code": code, "redirect_uri": REDIRECT, "code_verifier": session.pop("pkce"),
}, timeout=5)
r.raise_for_status()
return r.json() # id_token, access_token, refresh_token, expires_in, token_typeServer-to-server calls with no user use the client credentials grant on a confidential app client, which returns an access token with the custom scopes that client is allowed. Never embed a client secret in a browser or mobile app.
Three tokens with three jobs
| Token | What it says | Send it to | Lifetime |
|---|---|---|---|
| ID token | Who the user is: sub, email, name, groups; aud is the app client ID | Your own front end, and identity pools | 5 minutes to 1 day, set per app client |
| Access token | What the caller may do: scope, groups, client_id | Your APIs and resource servers | 5 minutes to 1 day, set per app client |
| Refresh token | Encrypted and opaque; only the user pool can read it | Only the user pool's token endpoint or API | Default 30 days; 60 minutes to 10 years |
The single most common mistake is sending the ID token to APIs. Authorise API calls with the access token, which carries scopes and is meant for that purpose; use the ID token to learn about the user in the client. A token's token_use claim says which kind it is.
Tokens grow when you add claims, so do not assume a fixed size. In browsers, keeping tokens in a server-side session behind an HTTP-only cookie (a backend-for-frontend) avoids exposing them to cross-site scripting; storing them in local storage is simpler and riskier.
Verifying tokens correctly
Your API must verify every token. Decoding a JWT is not verifying it. Cognito signs tokens with RS256 and publishes the public keys at https://cognito-idp.<region>.amazonaws.com/<userPoolId>/.well-known/jwks.json. Verification means: find the key matching the token's kid header, check the signature, check exp, check that iss is your pool, check token_use, and check the audience, which is aud in ID tokens but client_id in access tokens.
import jwt # PyJWT
from jwt import PyJWKClient
REGION, POOL, CLIENT_ID = "eu-west-1", "eu-west-1_EXAMPLE", "1example23456789"
ISSUER = f"https://cognito-idp.{REGION}.amazonaws.com/{POOL}"
jwks = PyJWKClient(f"{ISSUER}/.well-known/jwks.json") # caches keys by kid
def verify_access_token(token, required_scope):
key = jwks.get_signing_key_from_jwt(token).key
claims = jwt.decode(token, key, algorithms=["RS256"], issuer=ISSUER,
options={"verify_aud": False, "require": ["exp", "iss", "token_use"]})
if claims["token_use"] != "access":
raise PermissionError("not an access token")
if claims.get("client_id") != CLIENT_ID: # access tokens carry client_id, not aud
raise PermissionError("wrong app client")
if required_scope not in claims.get("scope", "").split():
raise PermissionError("missing scope")
return claims # sub, username, cognito:groups, scope, exp, ...Cache keys by kid and refetch when an unknown kid appears, because Cognito can rotate signing keys. In Node.js, AWS recommends the aws-jwt-verify library, which implements these checks. If you put API Gateway in front, a Cognito authorizer on a REST API or a JWT authorizer on an HTTP API can do the verification for you. For the HTTP API JWT authorizer, set the issuer to the pool URL and the audience to the app client ID; access tokens bound to a resource server carry that resource identifier as aud instead, so configure the audience accordingly.
Refresh, rotation and signing out
When the access token expires, the client uses the refresh token to get new ones, through the OAuth token endpoint with grant_type=refresh_token, the GetTokensFromRefreshToken API, or the older REFRESH_TOKEN_AUTH flow. Enable refresh token rotation on the app client: each refresh then returns a new refresh token and invalidates the old one, with an optional grace period of up to 60 seconds for retries. Rotation does not extend the session; every rotated token expires at the end of the original validity window. Rotation is incompatible with REFRESH_TOKEN_AUTH, so switch clients to GetTokensFromRefreshToken or the token endpoint first.
Signing out has layers. RevokeToken ends one refresh-token session; GlobalSignOut (with the user's access token) or AdminUserGlobalSignOut (with AWS credentials) ends all of them. Managed login also keeps a session cookie, valid for an hour, so after signing a user out, redirect them to the logout endpoint or they can sign straight back in. Most importantly, a signed access token verified offline stays valid until it expires even after revocation. If revocation must take effect quickly, keep access token lifetimes short.
Lambda triggers
User pools call your Lambda functions at defined points: pre sign-up, post confirmation, pre and post authentication, pre token generation, user migration, custom message, the three custom-challenge triggers, inbound federation, and custom email and SMS senders. Except for the custom senders, Cognito invokes them synchronously and the function must respond within 5 seconds, a limit you cannot change. A slow or failing trigger fails the user's sign-in, so keep triggers small, avoid slow network calls, and alarm on their errors and duration. Error messages you raise are shown to users, so log details to CloudWatch and return only safe text.
The pre token generation trigger is the most useful. With event version V1_0 it customises the ID token. With V2_0, available in Essentials and Plus, it can also add, override and suppress claims and scopes in access tokens; V3_0 extends that to machine-to-machine tokens.
# Pre token generation, event version V2_0: add a tenant claim and a scope to the access token.
def handler(event, context):
attrs = event["request"]["userAttributes"]
tenant = attrs.get("custom:tenant_id")
if not tenant:
raise Exception("Account is not assigned to a tenant") # text is shown to the user
event["response"]["claimsAndScopeOverrideDetails"] = {
"idTokenGeneration": {
"claimsToSuppress": ["phone_number"],
},
"accessTokenGeneration": {
"claimsToAddOrOverride": {"tenant_id": tenant},
"scopesToAdd": [f"tenant/{tenant}"] if " " not in tenant else [],
},
}
return event # must return within 5 seconds: no slow network calls hereThe migrate user trigger lets you move users from an old directory lazily: on a failed sign-in or password reset, Cognito calls your function, which checks the old system and returns the user's attributes. Use it if you are migrating in, because user passwords cannot be bulk-imported with their hashes. For Lambda itself, see the Lambda article.
Identity pools: AWS credentials for clients
When a mobile app must upload directly to S3 or publish to IoT, an identity pool turns a user pool ID token into temporary credentials. The recommended enhanced flow is two calls: GetId returns a stable identity ID, and GetCredentialsForIdentity returns credentials, valid for one hour, for the role the pool selects. The basic flow instead calls GetOpenIdToken and then STS AssumeRoleWithWebIdentity, which lets the client pick the role; AWS recommends leaving it off unless you need it.
Role selection can be a default authenticated role, rules on token claims (role-based access control), or session tags derived from claims for attribute-based access control, so one role can restrict each user to their own prefix with a policy condition such as ${aws:PrincipalTag/tenant}. IAM policy conditions explains the condition syntax. Treat the unauthenticated role as public: anyone can obtain its credentials.
Decisions that are hard to undo
- Sign-in identifiers. Whether users sign in with a username, or with email or phone as the username or an alias, is fixed when the pool is created. Choose carefully; changing it means a new pool and a migration.
- Custom attributes. They cannot be removed or renamed after creation, so keep them few and generic.
- Attribute write permissions. By default an app client can let users update their own attributes. If you put authorisation data such as
custom:tenant_idor a role in an attribute, remove write permission for it on every app client, or users can grant themselves access. - Region. A user pool lives in one Region. Plan multi-Region resilience explicitly rather than assuming replication.
- Tenancy model. One pool per tenant gives strong isolation and per-tenant settings but multiplies operations; a shared pool with a tenant attribute and claims is simpler but puts isolation in your code. Decide before you have many users.
Failure modes and operations
| Symptom | Cause | Fix |
|---|---|---|
| Sign-ins fail intermittently | Trigger exceeding 5 s or a cold dependency | Trim the trigger, cache, alarm on duration |
| API accepts tokens from another app | Audience or client_id not checked | Verify issuer, token_use and client |
| User escalates privileges | Writable authorisation attribute | Remove attribute write permission |
| Signed-out user still calls API | Offline verification cannot see revocation | Short access token lifetime |
| Verification emails not arriving | Default Cognito email sender used in production | Configure Amazon SES for the pool |
| Throttling during launch | Request-rate quotas on user pool APIs | Check quotas, cache tokens, request increases early |
Watch CloudWatch metrics for sign-in and token refresh successes and throttles, log administrative API calls with CloudTrail, and test sign-in end to end with a synthetic user. If you expose GraphQL, AppSync accepts user pool tokens natively. The trade-off against a third-party identity provider or a self-hosted server such as Keycloak is mainly control versus operations: Cognito is cheap at scale and integrates natively with AWS, but has less flexible customisation and some irreversible settings.
What to do next
- Decide whether you need an identity pool at all; if clients only call your APIs, use a user pool alone.
- Choose the feature plan from the features you need, and fix the sign-in identifier and tenancy model before creating the production pool.
- Create a public app client with the authorization code flow and PKCE, no secret, and exact callback URLs.
- Verify access tokens in your API, checking signature, expiry, issuer, token_use, client and scope, or delegate this to an API Gateway authorizer.
- Enable refresh token rotation, keep access tokens short, and implement sign-out including the logout endpoint.
- Remove user write permission on every attribute used for authorisation, configure SES for email, and alarm on trigger errors and duration.