Amazon Comprehend is AWS's managed natural language processing service. You send it UTF-8 text and it returns entities, key phrases, the dominant language, sentiment, syntax, personal data spans or toxicity scores, or the output of a classifier or entity recognizer you trained on your own labels. It needs no model hosting or GPUs for the built-in features, and its outputs are structured, scored and offset-addressed, which is exactly what downstream code needs.
Using it well is mostly about limits and operating modes. Document size limits differ by operation and are measured in bytes, synchronous calls are throttled dynamically, asynchronous jobs need an IAM role and S3 layout, and custom endpoints are billed while they run whether or not you send traffic. This article covers the feature set as it stands, the three processing modes, working Python for each, PII redaction, custom classification with endpoint capacity arithmetic, a worked triage pipeline, when a large language model is a better fit, and the failure modes. Limits quoted here are from the Comprehend developer guide at the time of writing; check the current quotas page before you design around them.
What Comprehend does today
| Capability | Synchronous API | Real-time document size limit |
|---|---|---|
| Entities, key phrases, dominant language | DetectEntities, DetectKeyPhrases, DetectDominantLanguage | 100 KB |
| Sentiment, targeted sentiment, syntax | DetectSentiment, DetectTargetedSentiment, DetectSyntax | 5 KB |
| PII location | DetectPiiEntities, ContainsPiiEntities | 100 KB |
| Toxicity (English) | DetectToxicContent | Up to 10 segments of 1 KB each |
| Custom classification | ClassifyDocument on an endpoint | 10 KB of plain text |
| Custom entities | DetectEntities with an endpoint ARN | 5 KB of plain text |
Each built-in capability also has an asynchronous Start*Job form for large corpora, and a BatchDetect* form for most of the single-document detections. PII redaction, which returns a copy of the text with spans masked, is available only through the asynchronous job. Toxicity detection currently supports English only, and PII detection is documented for English and Spanish.
The feature set has also shrunk. AWS stopped offering topic modeling, event detection and prompt safety classification to new customers from 30 April 2026. Accounts that used them in the previous 12 months keep access, and AWS points everyone else to Amazon Bedrock models for topics and events and Bedrock Guardrails for prompt safety. If you are designing something new, do not build on those three.
Three ways to call it
Single-document synchronous calls suit request-path work, such as checking a support message for PII before storing it. Latency is low and the result comes back in the response. Batch synchronous calls take up to 25 documents of at most 5 KB each and return a ResultList plus an ErrorList indexed by position. A batch can partly succeed, and code that ignores ErrorList silently drops documents. Asynchronous jobs read a prefix in S3, either one document per file or one per line, and write a compressed result archive to another prefix. They suit backfills and nightly runs, accept up to 1 MB per document for entities, key phrases, PII and language, and are limited to 10 active jobs per operation by default.
AWS describes synchronous throttling as dynamic: the service gradually raises the rate it accepts if capacity is available. Treat TooManyRequestsException and other throttling errors as normal, retry with exponential backoff and jitter, and cap your own concurrency rather than relying on the service to shed load gracefully.
Synchronous calls without surprises
The most common bug is the byte limit. A 5 KB limit on sentiment is 5,000 bytes of UTF-8, not 5,000 characters, so text in scripts that need two to four bytes per character fails much earlier than expected. Chunk by encoded size on sentence boundaries:
import re, time, random, boto3
from botocore.config import Config
from botocore.exceptions import ClientError
cmp = boto3.client("comprehend", config=Config(retries={"mode": "adaptive", "max_attempts": 8}))
def chunks(text, limit=4800): # stay under the 5,000-byte limit
buf = ""
for sent in re.split(r"(?<=[.!?])\s+", text):
cand = (buf + " " + sent).strip()
if len(cand.encode("utf-8")) <= limit:
buf = cand
continue
if buf:
yield buf
while len(sent.encode("utf-8")) > limit: # pathological sentence: hard cut
cut = sent.encode("utf-8")[:limit].decode("utf-8", "ignore")
yield cut
sent = sent[len(cut):]
buf = sent
if buf:
yield buf
def batch_sentiment(docs, lang="en"):
out = [None] * len(docs)
for start in range(0, len(docs), 25):
part = docs[start:start + 25]
resp = cmp.batch_detect_sentiment(TextList=part, LanguageCode=lang)
for r in resp["ResultList"]:
out[start + r["Index"]] = r["Sentiment"], r["SentimentScore"]
for e in resp["ErrorList"]: # partial failures are normal
out[start + e["Index"]] = ("ERROR", e["ErrorCode"])
return outThe SDK's adaptive retry mode handles throttling with client-side rate limiting. Aggregating chunk-level sentiment back to a document is your decision: averaging scores weighted by chunk length is a reasonable default, but for complaints a single strongly negative chunk is often the signal you care about, so keep the minimum as well.
PII redaction at scale
For stored data, the asynchronous PII job is the workhorse. With Mode='ONLY_REDACTION' it writes a redacted copy of every document; with ONLY_OFFSETS it writes span locations and types for your own redaction code.
job = cmp.start_pii_entities_detection_job(
JobName="tickets-2026-10-02",
LanguageCode="en",
Mode="ONLY_REDACTION",
RedactionConfig={
"PiiEntityTypes": ["NAME", "EMAIL", "PHONE", "ADDRESS",
"CREDIT_DEBIT_NUMBER", "BANK_ACCOUNT_NUMBER", "SSN"],
"MaskMode": "REPLACE_WITH_PII_ENTITY_TYPE",
},
InputDataConfig={"S3Uri": "s3://acme-nlp-in/tickets/2026-10-02/",
"InputFormat": "ONE_DOC_PER_LINE"},
OutputDataConfig={"S3Uri": "s3://acme-nlp-out/pii/",
"KmsKeyId": "alias/nlp-output"},
DataAccessRoleArn="arn:aws:iam::111122223333:role/ComprehendDataAccess",
ClientRequestToken="tickets-2026-10-02", # idempotent retries
)
while True:
d = cmp.describe_pii_entities_detection_job(JobId=job["JobId"])
status = d["PiiEntitiesDetectionJobProperties"]["JobStatus"]
if status in ("COMPLETED", "FAILED", "STOPPED"):
break
time.sleep(60)The data access role needs read access to the input prefix, write access to the output prefix, and use of the KMS key, with a trust policy that lets the Comprehend service principal assume it. Use a ClientRequestToken derived from the batch name so a retried orchestration step does not start a duplicate job. In production, replace the polling loop with an EventBridge rule or a Step Functions wait state.
Treat PII detection as a strong filter, not a guarantee. It is a statistical model with a score per span, it recognises a fixed list of types, and several government-ID types are country-specific. Combine it with deterministic checks for formats you know, such as your own account-number pattern, and sample the redacted output for misses.
Custom classifiers and endpoints
When the built-in labels do not match your domain, train a custom classifier with CreateDocumentClassifier or a custom entity recognizer with CreateEntityRecognizer. For a CSV-format multi-class classifier the guide requires at least 50 training documents per class; multi-label mode supports 2 to 100 labels. Custom entity recognizers support up to 25 entity types and need at least 25 annotations per type when trained from annotated plain text. Training reports precision, recall and F1 on a held-out split, and flywheels can manage retraining and version comparison as you add labelled data.
Serving a custom model in real time requires an endpoint, sized in inference units (IUs). Each IU provides up to 100 characters per second and 2 documents per second, an endpoint can have up to 50 IUs, and you pay for provisioned IUs for as long as the endpoint exists, idle or not. Asynchronous classification jobs need no endpoint and are usually the cheaper choice for anything that can wait.
Worked example: support ticket triage
A support team wants every incoming ticket classified into one of 14 queues and checked for personal data before it is stored. Volume is 40,000 tickets a day, averaging 1,200 characters, with a peak hour carrying three times the average rate.
- Capacity. Average throughput is 40,000 x 1,200 / 86,400, about 556 characters per second, which needs 6 IUs at the 100-characters-per-second rate. The document rate, about 0.46 per second, needs only one IU, so characters are the binding constraint. The peak is three times that, about 1,670 characters per second, or 17 IUs.
- Decide what must be real-time. Routing a ticket to a queue within seconds matters; redaction for the analytics copy does not. So classify in real time on an endpoint, and redact in an hourly asynchronous PII job over the analytics copy.
- Trim the input. ClassifyDocument rejects plain text over 10 KB rather than truncating it, so you must cut long tickets anyway, and the subject plus the first two paragraphs usually carry the intent. Truncating to 600 characters halves the IU requirement to about 9 at peak, and an offline evaluation on 2,000 labelled tickets shows whether accuracy holds.
- Route by confidence. Send tickets whose top class scores below a threshold, chosen from the evaluation set, to a human triage queue instead of guessing. Track the share that falls through; a rising share is the earliest sign of drift.
- Scale with the day. Use
UpdateEndpointwith a newDesiredInferenceUnitson a schedule, more in business hours and fewer at night, or Application Auto Scaling for the endpoint. Throttled calls fall back to a queue and retry.
The result is a predictable cost line, structured outputs your routing code can trust, and a fallback path that keeps uncertain tickets out of the wrong queue.
Comprehend or a large language model
A large language model on Amazon Bedrock can do all of these tasks from a prompt, often without labelled training data, and handles labels you invent tomorrow. Comprehend wins when you need stable, character-offset outputs for redaction, a fixed and auditable label set, predictable throughput and cost per character, and no prompt to maintain. LLMs win for open-ended extraction, labels that change often, and long or messy documents that need reasoning. Many systems use both: Comprehend for PII and coarse routing on every message, and an LLM for the smaller set that needs summarising or nuanced classification. If an LLM sees customer text, run PII redaction first.
Failure modes and trade-offs
- Byte limits mistaken for characters. Multi-byte text fails with a size error at a fraction of the expected length. Chunk on encoded bytes.
- Ignoring ErrorList. Batch calls partly succeed; dropped indices become silent data loss.
- Idle endpoints. An endpoint left running after a test bills every hour. Tag endpoints with an owner and delete or scale them down on a schedule.
- Wrong language code. Sending Spanish text with
enreturns poor results rather than an error. Detect the language first or route by a known locale. - Treating scores as probabilities you can compare across models. Calibrate thresholds per model on labelled data, and re-check them after retraining.
- Offset handling. Slicing redaction spans from a string that was normalised or re-encoded after detection shifts every offset. Redact against exactly the text you sent, and include non-ASCII cases in tests.
- Region assumptions. Comprehend is available in a subset of Regions. Check that your data's Region is on the list before designing around it.
Related reading: Amazon Textract for getting text out of scanned documents before Comprehend sees it, Amazon Bedrock for the LLM path, SageMaker if you outgrow managed models, and LLM deployment hardening for where PII filtering fits in a generative AI stack.
What to do next
- List the text you process and, for each stream, whether it needs a real-time answer or can be handled in batch.
- Check each operation's byte limit against your largest documents and add byte-based chunking where needed.
- Wrap batch calls so every
ErrorListentry is retried or recorded. - Run an asynchronous PII job over a sample, review the misses, and add deterministic checks for formats you own.
- If you need custom labels, gather at least 50 examples per class, train, and read precision and recall per class before deploying.
- Size any endpoint from characters per second at peak, schedule its capacity, and tag it with an owner.
- Avoid topic modeling, event detection and prompt safety for new designs; use Bedrock and Guardrails instead.