Microsoft Entra ID is the cloud identity provider behind Microsoft 365, Azure and many thousands of third-party applications. Until July 2023 it was called Azure Active Directory, and plenty of documentation, SDKs and PowerShell modules still use the old name. Whatever it is called, it answers two questions for every request: who is this, and what are they allowed to get a token for.
Despite the old name, it is not Active Directory in the cloud. It has no LDAP, no Kerberos for arbitrary servers, no organisational units and no Group Policy; it speaks OAuth 2.0, OpenID Connect and SAML over HTTPS. This article explains the object model that confuses most newcomers, how sign-ins become tokens, how to validate those tokens, how Conditional Access makes decisions, how workloads authenticate without secrets, and what goes wrong. Provider-neutral IAM design is covered in cloud IAM.
The tenant and its object model
A tenant is a dedicated instance of Entra ID for one organisation, identified by a GUID tenant ID and one or more verified domain names. Everything else lives inside a tenant: users, groups, devices, applications and policies. Azure subscriptions trust exactly one tenant for identity, and a user can be a member of one tenant and a guest (B2B collaboration) in others.
The part that trips everyone up is applications. An app registration is the global definition of an application: its client ID (also called application ID), redirect URIs, credentials, the scopes it exposes and the app roles it defines. A service principal, shown in the portal as an Enterprise Application, is the local instance of that application in a particular tenant. It is the object that gets assigned roles, that users consent to and that Conditional Access targets. A single-tenant line-of-business app has one registration and one service principal in the same tenant. A multi-tenant SaaS app has one registration in the vendor's tenant and a service principal in every customer tenant that has consented to it.
Groups are either security groups or Microsoft 365 groups, with assigned or dynamic membership. Dynamic groups compute membership from user attributes such as department, which makes joiner and leaver handling automatic but makes the attribute source a security-critical system.
Two permission planes, and scopes versus app roles
Entra ID has two separate authorisation systems that are easy to conflate. Entra roles such as Global Administrator, User Administrator and Application Administrator govern the directory itself: who can create users, reset passwords or register apps. Azure RBAC roles such as Owner, Contributor and Storage Blob Data Reader govern Azure resources, assigned at management group, subscription, resource group or resource scope. A Global Administrator has no access to a storage account by default, and a subscription Owner cannot reset anyone's password.
Inside your own applications there is a second split. Delegated permissions, or scopes, let an application act on behalf of a signed-in user; they appear in the access token's scp claim, and the effective access is the intersection of what the app was granted and what the user can do. Application permissions, exposed as app roles with allowed member type Application, let a daemon act as itself; they appear in the roles claim and require admin consent. App roles assigned to users and groups also appear in roles, and are usually a better way to express application authorisation than group membership.
How a sign-in becomes a token
Interactive applications use the OAuth 2.0 authorization code flow with PKCE. The app redirects the browser to https://login.microsoftonline.com/<tenant>/oauth2/v2.0/authorize with its client ID, redirect URI, requested scopes and a code challenge. Entra ID authenticates the user, applies Conditional Access, asks for consent if needed and redirects back with a one-time code. The app exchanges that code at the /oauth2/v2.0/token endpoint for an ID token (who the user is, for the app), an access token (for one resource API) and a refresh token.
Daemons use the client credentials flow: the service authenticates as its own service principal and asks for <resource>/.default, which means all application permissions already granted. With the MSAL library it looks like this:
import msal
app = msal.ConfidentialClientApplication(
client_id=CLIENT_ID,
authority="https://login.microsoftonline.com/" + TENANT_ID,
client_credential=CERT_OR_SECRET, # prefer a certificate; better still, a managed identity
)
result = app.acquire_token_for_client(scopes=["api://orders-api/.default"])
if "access_token" not in result:
raise RuntimeError(result.get("error_description"))
headers = {"Authorization": "Bearer " + result["access_token"]}MSAL caches tokens in memory and returns a cached token until it nears expiry, so calling it per request is cheap. Access tokens get a default lifetime chosen at random between 60 and 90 minutes, which spreads refresh load. Do not build logic that assumes an exact expiry.
Token anatomy and validation
Access tokens for your own APIs are JWTs signed with keys the tenant publishes at its OpenID Connect discovery document. An API must check the signature, the issuer (iss, which identifies the tenant and token version), the audience (aud, which must be your API), the expiry, and then authorise on scp or roles. Use oid together with tid as the stable user identifier, never the email or UPN, which can change and can be reassigned. Whether your API receives v1.0 or v2.0 tokens is set by the access token version in the API's app registration manifest, not by which endpoint the client called.
import jwt
from jwt import PyJWKClient
TENANT_ID = "..." # your tenant GUID
API_CLIENT_ID = "..." # the API's application (client) ID
ISSUER = "https://login.microsoftonline.com/" + TENANT_ID + "/v2.0"
jwks = PyJWKClient("https://login.microsoftonline.com/" + TENANT_ID + "/discovery/v2.0/keys")
def authorize(bearer, required_scope=None, required_role=None):
key = jwks.get_signing_key_from_jwt(bearer).key # cached, refetched on unknown kid
claims = jwt.decode(bearer, key, algorithms=["RS256"],
audience=API_CLIENT_ID, issuer=ISSUER)
scopes = claims.get("scp", "").split()
roles = claims.get("roles", [])
if required_scope and required_scope not in scopes:
raise PermissionError("missing scope " + required_scope)
if required_role and required_role not in roles:
raise PermissionError("missing role " + required_role)
return claims["tid"], claims["oid"]Two details bite in production. Group claims overflow: when a user belongs to more groups than fit (200 for JWTs), the token carries an overage indicator instead of the list, and your API must call Microsoft Graph or, better, rely on app roles. And tokens issued for Microsoft Graph are not meant to be validated by you; only validate tokens whose audience is your own API.
Conditional Access: the policy engine
Conditional Access (CA) is evaluated after the first authentication factor and before a token is issued. Each policy has assignments (which users, groups or workload identities, which target applications, and conditions such as location, device platform, client app type and sign-in or user risk) and controls: block, or grant only if the user completes MFA, uses a phishing-resistant authentication strength, signs in from a compliant or hybrid-joined device, or accepts terms. Session controls set sign-in frequency and restrict browser sessions. All matching policies apply, and every grant requirement must be satisfied; a block wins.
Build policies in report-only mode first and read the sign-in logs to see what would have happened. A sensible baseline requires MFA for all users, requires phishing-resistant methods for administrator roles, blocks legacy authentication protocols that cannot do MFA, and requires compliant devices for sensitive apps. Always exclude two emergency (break-glass) accounts from CA, protected by strong credentials kept offline and monitored for any use, so a policy mistake cannot lock everyone out. Conditional Access requires Entra ID P1 licensing; risk-based conditions and Privileged Identity Management require P2 or the ID Governance add-on. See zero trust architecture for how CA fits a broader design.
Workload identities without secrets
Client secrets on service principals are the most common credential leak in Azure estates: they are pasted into pipelines, expire unnoticed and outage production. Two mechanisms remove them. Managed identities are service principals whose credentials Azure manages for a resource such as a VM, App Service or function. System-assigned identities share the resource's lifecycle; user-assigned identities are standalone and can be attached to several resources. Code obtains tokens from the local instance metadata service, and the Azure SDKs do it for you:
# Raw call from inside an Azure VM (what the SDK does underneath)
curl -s -H "Metadata: true" \
"http://169.254.169.254/metadata/identity/oauth2/token?api-version=2018-02-01&resource=https://storage.azure.com/"
# Python SDK: works with managed identity in Azure and with developer credentials locally
from azure.identity import DefaultAzureCredential
from azure.storage.blob import BlobServiceClient
cred = DefaultAzureCredential()
blobs = BlobServiceClient("https://ordersdata.blob.core.windows.net", credential=cred)Workload identity federation covers workloads outside Azure. You add a federated identity credential to an app registration or user-assigned managed identity, naming an external issuer, a subject and an audience. The external platform issues its own OIDC token, and Entra ID exchanges it for an access token without any stored secret. For GitHub Actions, the issuer is https://token.actions.githubusercontent.com, the subject is something like repo:contoso/orders:ref:refs/heads/main or repo:contoso/orders:environment:production, and the audience is api://AzureADTokenExchange; the workflow needs id-token: write permission. Kubernetes clusters use the same mechanism with the cluster's service-account issuer, which is how AKS workload identity works; see managed Kubernetes. Scope subjects narrowly: a subject matching any branch lets any branch deploy to production.
Hybrid identity: syncing from on-premises Active Directory
Organisations with Windows Server Active Directory usually keep it as the source of truth and synchronise users and groups into Entra ID with Entra Connect Sync or the lighter, agent-based Entra Cloud Sync. Authentication then uses one of three methods. Password hash synchronisation copies a hash of the AD password hash to the cloud, so cloud sign-in keeps working when on-premises is down and leaked-credential detection works. Pass-through authentication has on-premises agents validate passwords live, so cloud sign-in depends on those agents and on the domain controllers. Federation with AD FS hands authentication to an on-premises farm, the most complex and least resilient option. Most organisations should use password hash sync, even alongside another method, as a fallback.
Sync is not instantaneous: a disabled AD account keeps working in the cloud until the next sync cycle and until existing tokens expire, unless continuous access evaluation revokes sessions for supporting services. For urgent leavers, disable and revoke sessions in Entra ID directly as well.
Worked example: securing an orders API end to end
A team runs an orders API on Azure App Service, a single-page web front end and a nightly reconciliation job in GitHub Actions. The setup: register the API, set its access token version to 2, expose a delegated scope Orders.Read and define two app roles, Orders.Admin for users and Orders.Reconcile for applications. Register the SPA as a public client using authorization code with PKCE, requesting api://orders-api/Orders.Read. Assign the Orders.Admin role to a security group of support leads.
For the batch job, create a user-assigned managed identity with a federated credential whose subject is the repository's production environment, grant it the Orders.Reconcile app role on the API's service principal, and grant it Storage Blob Data Reader on one container through Azure RBAC. The API itself uses its own managed identity to reach the database. Conditional Access requires MFA for the SPA and a compliant device for anyone holding the admin role. The result has no stored secrets anywhere, authorisation lives in app roles rather than group names, and every access path is visible in the sign-in logs.
Failure modes
- Expired client secret. A secret created with a one- or two-year lifetime expires and a production integration stops. Inventory credentials with Microsoft Graph and alert well before expiry, or move to managed identities and federation.
- Audience mismatch. The client requests a token for Microsoft Graph and sends it to your API, or the API expects the App ID URI while v2 tokens carry the client ID. Log the rejected claim, not just 401.
- Consent confusion. Users see an approval-required prompt because the app asks for a permission that needs admin consent. Grant consent deliberately, and restrict user consent to verified publishers and low-risk permissions to limit illicit consent grants.
- Lockout by policy. A CA policy blocks all administrators. This is what break-glass accounts and report-only mode prevent.
- Over-privileged service principals. An automation identity holds Application.ReadWrite.All or a directory role it no longer needs. Review app permissions like human access.
- Group overage. Authorisation silently fails for the users with the most groups, usually administrators.
Operational guidance
Stream sign-in logs, audit logs and service principal sign-ins to your log platform and alert on break-glass use, new credentials added to service principals, consent grants and role assignments. Use Privileged Identity Management so administrator roles are eligible and activated just in time with approval and justification, and run access reviews on privileged roles and guest users. Keep app registrations owned by named teams, managed as code where possible, and store any unavoidable secrets in a vault; see cloud secrets management.
What to do next
- Create two break-glass accounts excluded from Conditional Access, with strong offline credentials and alerts on use.
- Export every app registration and service principal credential with its expiry, and plan the migration of each secret to a managed identity or federated credential.
- Move CI/CD authentication to workload identity federation with subjects scoped to protected branches or environments.
- Deploy baseline Conditional Access policies in report-only mode, review a week of sign-in logs, then enable them.
- Review your APIs: validate issuer, audience and signature, authorise on scopes and app roles, and key users on oid plus tid.
- Put administrator roles behind just-in-time activation and schedule quarterly access reviews.