Amazon Macie answers a question that most S3 estates cannot answer from configuration alone: which buckets actually hold personal data, credentials or financial records, and which of those are exposed. It does two jobs that are easy to confuse. It continuously monitors bucket posture (public access, sharing, replication, encryption defaults) and raises policy findings when a change weakens it. It also reads object content, through daily sampled automated discovery or through jobs you define, and raises sensitive data findings when detectors match.
This article covers both halves, how to build and test a custom identifier, how to route findings, and the traps that leave buckets silently unscanned. Limits were checked against the Macie documentation on 3 October 2026; prices are left out because they change.
What Macie watches and what it ignores
Macie is scoped to Amazon S3 general purpose buckets in one Region. It does not inspect databases, volumes or data in transit, and it is not a threat detector; suspicious API activity belongs to GuardDuty. What Macie adds is content awareness. GuardDuty can tell you an unusual principal read a bucket; Macie can tell you that bucket holds passport numbers.
When you enable Macie it creates a service-linked role and builds an inventory of your buckets and their objects, refreshed daily. In an organisation, a delegated administrator sees inventory, findings and settings for every member account in that Region. Everything is per Region: a bucket in a Region where Macie is off is invisible to it.
Policy findings come from the posture half. Macie raises one only when a change happens after Macie was enabled, so a bucket that was already public on day one does not produce Policy:IAMUser/S3BucketPublic. Repeated occurrences update the existing finding and increment its count. The five types are:
| Finding type | Raised when |
|---|---|
Policy:IAMUser/S3BlockPublicAccessDisabled | All bucket-level block public access settings were turned off. |
Policy:IAMUser/S3BucketPublic | An ACL or bucket policy now allows anonymous users or all authenticated AWS identities. |
Policy:IAMUser/S3BucketSharedExternally | An ACL or policy shares the bucket with an account outside your organisation. |
Policy:IAMUser/S3BucketReplicatedExternally | Replication now targets a bucket owned by an external account. |
Policy:IAMUser/S3BucketSharedWithCloudFront | The policy grants a CloudFront origin access identity or origin access control. |
A sixth type, S3BucketEncryptionDisabled, now only means default encryption was reset to SSE-S3. The documentation warns that S3BucketSharedExternally can be a false positive when Macie cannot fully evaluate conditions such as aws:PrincipalOrgID or aws:SourceVpce in a bucket policy, so treat it as a prompt to read the policy, not a verdict.
Architecture: posture, content and results
Sensitive data findings come from the content half and always name a single object. Unlike policy findings, every detection is a new finding, even for an object reported yesterday. The five types are SensitiveData:S3Object/Credentials (secret access keys, private keys), /Financial (card and bank account numbers), /Personal (PII and health identifiers), /CustomIdentifier (your own detectors) and /Multiple when an object contains more than one category.
Findings are kept in Macie for 90 days. So are sensitive data discovery results, the per-object analysis records, but those are not visible in the console or API at all until you configure a repository: an S3 bucket plus a customer managed, symmetric KMS key in the same Region. Macie then writes gzip-compressed JSON Lines files there, including records for objects it analysed and found nothing in, and objects it could not analyse. That is your only durable proof of coverage, so configure the repository first; Macie back-fills the previous 90 days when you save it.
Detectors: managed, custom and allow lists
Detection runs three kinds of criteria over each object's extracted text. Managed data identifiers are AWS-maintained detectors for credentials, financial data and country-specific personal identifiers; they combine patterns with checks such as keyword proximity. Custom data identifiers are yours: a regular expression plus optional keywords, ignore words and a proximity distance. Allow lists subtract: text that matches a detector and also matches an allow-list entry is not reported.
The custom identifier rules are specific, and most false positives come from ignoring them:
- The regex can be up to 512 characters and uses a PCRE subset with no backreferences, capturing groups, lookahead or lookbehind, and no global flags; use
(?i)for case-insensitive parts. Large bounded repeats such as\d{100,1000}do not compile.^and$anchor to the start and end of the file, not the line. - Up to 50 keywords of 3 to 90 characters, matched case-insensitively. A keyword must precede the match within the maximum match distance, 1 to 300 characters and 50 by default, measured from the end of the keyword to the end of the match. In CSV, TSV and Excel files a keyword in the column name also counts; in JSON, Parquet and Avro a keyword in the field path counts.
- Up to 10 ignore words of 4 to 90 characters, matched case-sensitively; a match that contains one is dropped.
- Findings default to Medium severity regardless of count. You can instead set occurrence thresholds for Low, Medium and High, in ascending order, and anything below the lowest threshold produces no finding at all.
- A custom identifier cannot be edited after creation, so that historical findings stay explainable. Version the name, for example
employee-id-v2, and test before you create.
Allow lists come in two forms: a plain-text file in an S3 bucket you own in the same Region (1 to 100,000 entries, each 1 to 90 characters, at most 35 MB, case-insensitive exact matches), or a single regex of up to 512 characters stored in Macie. In structured files an entry only suppresses text held entirely in one cell or field, so the name "Akua Mansa" is suppressed in a single Name column but not when it is split across First Name and Last Name. If Macie cannot read or parse a file list at the start of a job or daily cycle, it analyses without it; check list status after every change.
Worked example: an employee ID detector
Suppose employee identifiers look like EMP-204817 and appear in HR exports, tickets and contract PDFs. A bare regex will also match the template value EMP-000000 and part numbers that share the prefix. Build the criteria in three steps and test each with TestCustomDataIdentifier, which accepts up to 1,000 characters of sample text and returns a single matchCount:
aws macie2 test-custom-data-identifier \
--regex "EMP-[0-9]{6}" \
--keywords "employee id" "emp no" "staff number" \
--ignore-words "EMP-000000" \
--maximum-match-distance 30 \
--sample-text "Employee ID: EMP-204817, cost centre 4410. Template row: EMP-000000. Part EMP-551200 ships Friday."EMP-204817 ends 12 characters after the keyword "Employee ID", inside the 30-character distance, so it counts. EMP-000000 contains an ignore word and is dropped. EMP-551200 has no keyword within 30 characters before it, so it is excluded. The expected matchCount is therefore 1; if you get a different number, your mental model of the rules is wrong. Then add counter-examples from real data and choose the distance just large enough for your formats. For CSV exports whose header is employee_id, add that as a keyword too, because column names count.
Finally set severity thresholds, for example Low from 1 occurrence, Medium from 50 and High from 1,000, so one ID in a ticket is not paged like a full HR export.
Automated discovery versus jobs
Content inspection runs in one of two modes, and most estates need both.
| Automated sensitive data discovery | Sensitive data discovery job | |
|---|---|---|
| Selection | Macie samples representative objects daily, grouped by bucket, prefix, storage class, extension and age, breadth first across buckets | Explicit buckets (up to 1,000) or runtime bucket criteria, plus object criteria (prefix, extension, size, tags, last modified) |
| Depth | Incremental; new and changed objects are prioritised; unchanged analysed objects are not re-read | Sampling depth is a percentage of objects; each chosen object is read in full |
| Schedule | Continuous daily cycle; first results within about 48 hours | One-time, or daily, weekly or monthly; periodic runs only analyse objects created or changed since the last run |
| Output | Findings, results, and a per-bucket sensitivity score | Findings and results; does not change sensitivity scores |
| Use for | Estate-wide map of where sensitive data lives | Proof for a specific bucket, deep scans, audits, remediation verification |
Automated discovery assigns each bucket a sensitivity score. A new bucket starts at 50, labelled Not yet analyzed; empty buckets get 1. Scores from 1 to 49 mean Not sensitive (a high number in that range means Macie has seen little of the bucket), 51 to 99 mean Sensitive, 100 is only ever a manual override, and -1 means every attempt to analyse the bucket's objects failed. A stuck 50 or a -1 usually means Macie cannot read the data, not that it is clean.
Routing findings so someone acts
Macie publishes every new finding to EventBridge as soon as it is processed, with source aws.macie and detail-type Macie Finding. Updates to existing policy findings are batched and published every 15 minutes by default; the interval is configurable. Security Hub is an optional extra destination, and by default only policy findings go there. Suppressed findings are not published anywhere. A rule for high-severity content findings looks like this:
{
"source": ["aws.macie"],
"detail-type": ["Macie Finding"],
"detail": {
"category": ["CLASSIFICATION"],
"severity": { "description": ["High"] }
}
}Route it to a function that summarises the finding. The fields below appear in the documented event schema; policy findings have classificationDetails set to null and no object.
import json
import os
import boto3
sns = boto3.client("sns")
TOPIC_ARN = os.environ["TOPIC_ARN"]
def handler(event, context):
d = event["detail"]
if d.get("sample") or d.get("archived"):
return {"skipped": True}
bucket = d["resourcesAffected"]["s3Bucket"]
obj = d["resourcesAffected"].get("s3Object") or {}
cls = d.get("classificationDetails") or {}
counts = {}
for group in (cls.get("result") or {}).get("sensitiveData", []):
for det in group.get("detections", []):
counts[det["type"]] = counts.get(det["type"], 0) + det["count"]
public = bucket.get("publicAccess", {}).get("effectivePermission") == "PUBLIC"
msg = {
"finding_id": d["id"],
"type": d["type"],
"severity": d["severity"]["description"],
"bucket": bucket["name"],
"key": obj.get("key"),
"bucket_is_public": public,
"detections": counts,
"result_record": cls.get("detailedResultsLocation"),
}
subject = ("PUBLIC " if public else "") + d["type"]
sns.publish(TopicArn=TOPIC_ARN, Subject=subject[:100], Message=json.dumps(msg, indent=2))
return msgNotice what the message carries: counts by type and a pointer to the result record, never the sensitive values. Restrict Macie's retrieve-samples feature to a small investigation role, and archive every event, because the console forgets after 90 days. A content finding in a bucket whose effectivePermission is PUBLIC deserves a page; the rest can go to a daily triage queue.
When Macie cannot read your data
The commonest way for Macie to be enabled and useless is that it cannot read the objects. Three causes cover most cases:
- Restrictive bucket policies. A deny-all-except-these-roles policy also denies the Macie service-linked role. Macie then shows only partial bucket details and a score stuck at 50. Add an exception for the service-linked role's ARN in the deny statement.
- Customer managed KMS keys. Objects encrypted with SSE-KMS under a customer managed key are analysed only if the key policy allows the Macie service-linked role to decrypt. Objects using SSE-C or client-side encryption can never be analysed. See S3 encryption options and KMS key policies for the mechanics.
- Unclassifiable objects. Only supported storage classes and file formats are eligible, and size quotas apply to compressed and archive files.
Macie reports these as coverage data, broken down by reason. Review buckets with score 50 or -1 weekly and treat a rising count as a platform defect.
Failure modes
Failure modes worth naming:
- No repository configured. Symptom: an auditor asks which objects were scanned and found clean, and there is no answer older than 90 days. Fix it on day one.
- Assuming posture findings cover existing exposure. Policy findings fire on changes after enablement. Query the bucket inventory for public and externally shared buckets once, at enablement, instead.
- Jobs that never re-scan. A periodic job only looks at new or changed objects. If you add a custom identifier, existing objects are not re-read unless you create a new one-time job.
- Logging buckets in automated discovery. They consume sampling budget and add noise; exclude them (up to 1,000 exclusions).
Trade-offs and limits
Macie's cost scales with the number of buckets monitored, the objects automated discovery tracks and the bytes inspected, so the main trade-off is coverage against spend. A reasonable pattern is automated discovery everywhere except logging buckets, plus periodic jobs on the few Sensitive buckets that feed analytics or third parties. Macie is not real-time DLP: a new file may not be sampled for days. To block sensitive uploads, inspect at the write path and use Macie as the independent audit. Data copied into a database or vector store needs its own classification.
What to do next
- Enable Macie through a delegated administrator in every Region you use, and configure the discovery-results repository with a customer managed key before anything else.
- Export the bucket inventory once and list every public, externally shared or externally replicated bucket; do not wait for policy findings.
- Enable automated discovery, exclude logging buckets, and after 48 hours list every bucket with score 50 or -1 and fix bucket and key policies until coverage is complete.
- Write your first custom identifier with keywords and ignore words, test it with
test-custom-data-identifieragainst real positive and negative samples, then create it with a versioned name. - Create EventBridge rules for High content findings and for public-bucket policy findings, route them through a summarising function like the one above, and archive every event beyond 90 days.
- Read EventBridge in depth for rule design and dead-letter handling, and IAM to restrict who can retrieve sensitive data samples.