Access accumulates. People change teams and keep their old permissions, a contractor's role outlives the contract, an incident grants a broad role that nobody removes, and a service account created for a migration two years ago still holds write access to production. Each grant was reasonable when made. Together they form the attack surface that turns one phished credential into a breach.
An IAM review is the control that pushes back: a recurring decision, by someone accountable, about whether each identity should still hold each permission, followed by actual removal. Most organisations run reviews; far fewer get access removed by them, because reviewers face thousands of line items they cannot evaluate and approve them all. This article designs a review system around evidence instead: collect every entitlement, join it with real usage, rank by risk, ask small specific questions, revoke safely, and keep proof. The policy model itself is covered in IAM architecture; this page is about keeping the result of that model honest over time.
What a review must answer
Every review question has the same shape: should principal P still hold entitlement E on resource R, granted via path G? The path matters. A user may hold database admin because of a direct grant, because they are in a group that holds it, or because a role they can assume holds it. Revoking the wrong edge either does nothing or removes access from forty other people.
A useful system therefore answers four things for each edge: who holds it (and who is accountable for that identity, such as a manager or a service owner), how they got it, whether they used it recently, and how much damage it could do. Without the second you cannot revoke precisely; without the third reviewers guess; without the fourth they spend equal time on read access to a wiki and on organisation-wide administrator.
Collecting entitlements into a graph
Each source system has its own model: identity provider groups, cloud role bindings and policies, SaaS application roles, database grants, Kubernetes RBAC bindings, and code-hosting team memberships. Collectors pull each one on a schedule, read-only, and normalise it into a small set of node and edge types. Keep raw snapshots too, because when an auditor asks what access looked like on a given date you want the original data, not your interpretation of it.
-- nodes
CREATE TABLE principal (id TEXT PRIMARY KEY, kind TEXT, -- human | service | group | role
source TEXT, owner_id TEXT, hr_status TEXT);
CREATE TABLE entitlement (id TEXT PRIMARY KEY, system TEXT, -- aws:123456789012, postgres:billing
resource TEXT, permission TEXT, -- s3://invoices, write
sensitivity INT); -- 1 low .. 4 critical
-- edges, one row per grant path, valid-time so history is queryable
CREATE TABLE grant_edge (from_id TEXT, to_id TEXT, -- principal->group, group->role, role->entitlement
via TEXT, managed_by TEXT, -- 'terraform:repo/path' | 'console' | 'unknown'
valid_from TIMESTAMPTZ, valid_to TIMESTAMPTZ);
CREATE TABLE usage_fact (principal_id TEXT, entitlement_id TEXT,
last_used TIMESTAMPTZ, source TEXT, observed_at TIMESTAMPTZ);Effective access is the transitive closure from a human or service principal through groups and roles to entitlements. Compute it in the pipeline, and store the path, not only the end result, because the path is what a revocation edits. Cloud policy languages add conditions, deny statements and boundaries, so a naive closure overstates access; where the provider offers a policy simulator or an analyzer, use it to confirm effective permissions for high-sensitivity edges rather than reimplementing policy evaluation yourself.
Identity resolution is the unglamorous core. The same person appears as an email in the identity provider, a username in a database and a numeric id in a SaaS tool. Join on an immutable employee id from the HR system wherever possible, and route unmatched accounts into their own queue: an account nobody can attribute to a person or an owning service is the highest-value finding the system will produce.
Joining usage evidence
Reviewers cannot judge whether access is needed, but logs can say whether it was used. Sources include cloud audit trails, the last-accessed data that major cloud providers expose for roles and services, identity provider sign-in logs for application assignments, and database audit logs. Normalise them into usage_fact rows with the finest granularity the source supports, and record the granularity, because "used the storage service" is weaker evidence than "wrote to this bucket".
from datetime import datetime, timedelta, timezone
def classify(edge, usage, now=None, lookback_days=90):
now = now or datetime.now(timezone.utc)
if edge.principal.hr_status == "terminated":
return "revoke_now" # leaver: no review needed
if edge.age < timedelta(days=14):
return "too_new" # no evidence window yet
last = usage.get((edge.principal.id, edge.entitlement.id))
if edge.entitlement.periodic: # quarter-end jobs, DR drills, break-glass
lookback_days = max(lookback_days, 400)
if last is None or now - last > timedelta(days=lookback_days):
return "unused"
return "used"Usage evidence has three traps. First, absence of evidence is not evidence of absence when a log source has gaps or coarse granularity, so mark entitlements whose source cannot report usage as unknown rather than unused. Second, some legitimate access is rare: annual audits, disaster-recovery drills and break-glass roles must not be removed for being idle; tag them and give them a longer window or a different control. Third, "used" does not mean "needed" — an engineer who still queries the old team's database out of habit holds access they should lose.
Risk scoring decides what humans see
A reviewer can make perhaps a few hundred careful decisions in a sitting. Spend them where the risk is. A simple, explainable score beats a clever opaque one, because reviewers and auditors must understand why an item was raised.
def risk(edge, usage_class):
s = edge.entitlement.sensitivity * 10 # 10..40
s += {"unused": 25, "unknown": 10, "used": 0}[usage_class]
if edge.principal.kind == "service" and edge.principal.owner_id is None:
s += 30 # orphaned workload identity
if edge.managed_by == "console":
s += 15 # outside infrastructure-as-code
if edge.via == "direct" and edge.principal.kind == "human":
s += 5 # bypasses group-based model
if edge.principal.team_changed_since(edge.valid_from):
s += 20 # mover kept old access
return sItems above a high threshold go to a reviewer with the evidence attached. Low-risk used access can be certified in bulk or auto-approved with a recorded rule. Some decisions need no human at all: access held by terminated identities is removed automatically, and unused low-sensitivity access can be removed with notice. Tune the weights by looking at what reviewers actually revoke; if they revoke 40 percent of console-created grants and 2 percent of IaC-managed ones, the weight is earning its place.
Campaigns and event-driven reviews
Periodic campaigns, often quarterly for sensitive systems and annually for the rest, satisfy most compliance frameworks and catch slow drift. Assign each item to the person able to judge it: the manager for a human's business need, the resource owner for whether that role should exist at all, and the service owner for a workload identity. Give each campaign a deadline and a default outcome for items nobody decides. Defaulting to revoke for high-risk unused access is the strongest incentive to respond; defaulting to keep makes the campaign ceremonial.
Event-driven micro-reviews catch problems while context is fresh. When HR records a team change, ask the new manager within days which of the old team's access the person still needs, rather than waiting for the quarter. When the drift detector sees a console-created grant on a production role, open a review for that edge immediately. When access goes unused for the lookback window, ask the holder whether to keep it. Each of these is one specific question, which people answer accurately, instead of a spreadsheet of three thousand rows, which they do not.
Detecting rubber-stamping
A review whose approval rate is 100 percent and whose median decision time is two seconds has removed nothing and proved nothing. Measure the reviewers as well as the access. Useful signals are approval rate per reviewer, time per decision, the share of decisions made in one bulk action, and the rate at which approved-but-unused access is later revoked by another control.
SELECT reviewer_id,
count(*) AS decisions,
avg((decision = 'keep')::int) AS keep_rate,
percentile_cont(0.5) WITHIN GROUP (ORDER BY seconds_on_item) AS median_seconds,
avg((risk >= 60 AND usage_class = 'unused' AND decision = 'keep')::int) AS kept_risky_unused
FROM review_decision
WHERE campaign_id = $1
GROUP BY reviewer_id
HAVING count(*) >= 50
ORDER BY kept_risky_unused DESC;Feed the result back: require a written justification for keeping high-risk unused access, sample a fraction of decisions for a second reviewer, and redesign campaigns that produce universal approval, usually by cutting what reviewers are shown.
Staged, reversible revocation
A revocation that breaks a nightly batch job at 2 a.m. teaches the organisation to stop revoking. Design the executor to remove access in stages and to be reversible. First notify the holder and owner with the date. Then disable rather than delete where the system supports it, for example by detaching a policy or removing a group membership while keeping the definition. Watch for denied-access errors from that principal for a grace period, and restore automatically through the same pipeline if an owner confirms the access was needed. Only after the grace period delete the dormant artefacts.
Remember propagation. Removing a role binding does not end existing sessions, cached tokens or long-lived keys: issued credentials often remain valid until they expire. For leavers, revoke sessions and rotate any shared secrets the person could read as part of the same workflow. For grants managed as code, the executor should open a pull request against the repository instead of editing the live system, or the next apply will quietly restore the access it just removed. Give the executor remove-only permissions so a bug or a compromise of the review system cannot grant anything.
Drift against infrastructure-as-code
When IAM is managed as code, the review question changes from should this exist? to does reality match what was reviewed in the repository? Drift detection compares the collected live state with the desired state rendered from the code. Grants present live but absent from code were created by hand, often during an incident; grants in code but absent live mean someone removed them without updating the repository. Both should open an event-driven review. Code-managed grants still need periodic review, because a merged pull request two years ago is not a current business justification, but drift detection makes the code a trustworthy source of truth between campaigns.
Worked example: the billing database
The pipeline collects 1,180 edges touching the production billing database. Thirty-one belong to identities the HR feed marks as terminated; they are revoked automatically and sessions are killed. A service account svc-migrate-2024 has no owner and no use in 300 days; it scores 85 and goes to the database owner, who confirms it was a one-off migration and revokes it. Fourteen engineers moved teams last quarter and kept write access; their new managers are asked one question each and revoke twelve. A break-glass role unused for a year is tagged periodic and kept, with its sessions recorded under privileged access management. The remaining used, IaC-managed read grants are certified in bulk under a recorded rule. Humans make about 40 decisions instead of 1,180, and 57 grants disappear.
Failure modes
| Failure | What goes wrong | Defence |
|---|---|---|
| Spreadsheet campaigns | Reviewers approve everything | Risk ranking, evidence on each item, small specific questions |
| Revoking the wrong edge | Access survives or many people lose it | Store grant paths; revoke the path, not the end permission |
| Treating missing logs as unused | Needed access removed, trust in the program lost | Track source granularity; mark unknown separately |
| Idle-based removal of rare access | DR or break-glass fails during an incident | Tag periodic roles; longer windows or separate controls |
| Live edits to IaC-managed grants | Next apply restores revoked access | Executor opens pull requests for code-managed edges |
| Sessions outlive revocation | Leaver keeps access until token expiry | Revoke sessions and rotate shared secrets in the same workflow |
| Over-privileged review tooling | Compromise of the review system grants access | Read-only collectors, remove-only executor |
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Review cadence | Quarterly campaigns: simple, auditor-friendly, stale context | Event-driven: accurate and timely, needs reliable HR and drift feeds |
| Default for undecided items | Keep: no breakage, ceremonial control | Revoke high-risk unused: real reduction, needs staged rollback |
| Build vs buy | Identity governance product: connectors and workflow ready | In-house pipeline: fits your systems and IaC, you own the connectors |
| Usage lookback | Short (30 days): removes more, breaks rare jobs | Long (90+ days): safer, slower reduction |
Whichever you pick, layer it on a sound authorization model; RBAC and ABAC covers how role and attribute design determines how reviewable your grants are in the first place.
What to do next
- Inventory every system that grants access, including SaaS tools and databases, and write a read-only collector for the riskiest three first.
- Normalise into principals, entitlements and grant edges with paths and valid-time history; keep raw snapshots.
- Resolve identities on an immutable HR id and route unattributable accounts to an owner-assignment queue.
- Join usage evidence with its granularity, and tag periodic and break-glass roles before any idle-based removal.
- Automate leaver revocation including sessions, then add a simple, explainable risk score.
- Replace spreadsheet campaigns with risk-ranked items and event-driven questions on team changes and drift.
- Measure reviewer behaviour and require justification to keep high-risk unused access.
- Build a remove-only, staged executor that opens pull requests for code-managed grants, and store every decision as evidence.