Amazon Textract is a managed AWS service that reads documents. Give it a scanned PDF, a photo of a receipt or a multi-page TIFF and it returns the text it found, where on the page each word sits, and, if you ask, the structure: key-value pairs from forms, rows and columns from tables, answers to natural-language questions about the page, signatures, and layout elements such as titles and paragraphs. It is OCR plus document understanding, delivered as an API with no model to host.
What it returns is not a tidy JSON record of your invoice. It is a graph of Block objects linked by IDs, with confidence scores you must interpret. This page explains that model from first principles, maps the API families and their hard limits (checked against the AWS documentation on 2026-10-01), walks the graph in code, builds an idempotent asynchronous pipeline, and covers validation, review, failure modes and cost. The worked example extracts invoices end to end.
What Textract does, and what it does not
Textract performs three layers of work. Detection finds text: lines and words, printed or handwritten, with a bounding box and polygon for each. Analysis relates the text: which words form a form key and which form its value, which cells form a table, which checkbox is ticked. Specialised analysis understands specific document families: invoices and receipts through the expense API, US identity documents through the ID API, and mortgage packages through the lending API.
It does not validate business meaning: it will return a total of 1,250.00 with 99% confidence when the line items add to 1,205.00, because the page is wrong or a smudged digit was misread confidently. Documented language support is English, French, German, Italian, Portuguese and Spanish, with handwriting and queries English only, and vertical text is unsupported. Check those boundaries against your documents first.
The API families and their hard limits
| Need | Synchronous | Asynchronous |
|---|---|---|
| Text only | DetectDocumentText | StartDocumentTextDetection, GetDocumentTextDetection |
| Forms, tables, queries, signatures, layout | AnalyzeDocument | StartDocumentAnalysis, GetDocumentAnalysis |
| Invoices and receipts | AnalyzeExpense | StartExpenseAnalysis, GetExpenseAnalysis |
| US passports and driver's licences | AnalyzeID | none |
| Mortgage document packages | none | StartLendingAnalysis, GetLendingAnalysis |
The general analysis call takes a FeatureTypes list whose valid values are TABLES, FORMS, QUERIES, SIGNATURES and LAYOUT. Lines and words are always returned regardless of the features you request. Synchronous calls accept bytes or an S3 object and return results in the response; asynchronous calls read from S3 only, return a JobId, and deliver results through a separate paginated Get call.
The hard limits, which cannot be raised, shape the design. Synchronous operations accept JPEG, PNG, PDF and TIFF up to 10 MB, and PDF or TIFF input must be a single page. Asynchronous operations accept JPEG and PNG up to 10 MB and PDF or TIFF up to 500 MB and 3,000 pages. Password-protected PDFs and XFA-based PDFs are rejected. Images may be at most 10,000 pixels on a side, and text must be at least 15 pixels high to be detected, which the documentation equates to 8 point type at 150 DPI. Queries are limited to 15 per page synchronously and 30 per page asynchronously. A JobId is valid for 7 days. The practical rule: multi-page documents mean the asynchronous API, and low-resolution scans are a data quality problem you fix before calling Textract, not after.
The Block model
Every response is a flat list of Block objects. Each has an Id, a BlockType, a Page, a Geometry with a bounding box and polygon in coordinates normalised to the page from 0 to 1, usually a Confidence from 0 to 100, and a list of Relationships, each a type plus a list of IDs. Structure lives entirely in those relationships.
A PAGE block has CHILD links to its LINE blocks, and each line links to its WORD blocks. Forms appear as pairs of KEY_VALUE_SET blocks: the one whose EntityTypes contains KEY has a VALUE relationship to the other, and both have CHILD links to the words that make them up. Checkboxes are SELECTION_ELEMENT blocks with a SelectionStatus. Tables are a TABLE block with children of type CELL, each carrying a one-based RowIndex and ColumnIndex plus spans, with MERGED_CELL blocks for merged regions. A QUERY block holds your question and alias and links through an ANSWER relationship to a QUERY_RESULT block with the answer text and its confidence.
def index_blocks(blocks):
return {b["Id"]: b for b in blocks}
def child_text(block, by_id):
words = []
for rel in block.get("Relationships", []):
if rel["Type"] == "CHILD":
for cid in rel["Ids"]:
child = by_id[cid]
if child["BlockType"] == "WORD":
words.append(child["Text"])
elif child["BlockType"] == "SELECTION_ELEMENT":
words.append("[x]" if child["SelectionStatus"] == "SELECTED" else "[ ]")
return " ".join(words)
def key_values(blocks):
"""FORMS output: KEY blocks point at VALUE blocks; both point at their WORD children."""
by_id = index_blocks(blocks)
out = []
for b in blocks:
if b["BlockType"] == "KEY_VALUE_SET" and "KEY" in b.get("EntityTypes", []):
value_ids = [i for r in b.get("Relationships", []) if r["Type"] == "VALUE" for i in r["Ids"]]
value = " ".join(child_text(by_id[v], by_id) for v in value_ids)
out.append({"key": child_text(b, by_id), "value": value,
"key_conf": b["Confidence"], "page": b.get("Page", 1)})
return out
def tables(blocks):
"""TABLES output: TABLE -> CELL (RowIndex, ColumnIndex, 1-based) -> WORD."""
by_id = index_blocks(blocks)
result = []
for t in (b for b in blocks if b["BlockType"] == "TABLE"):
grid = {}
for rel in t.get("Relationships", []):
if rel["Type"] != "CHILD":
continue
for cid in rel["Ids"]:
cell = by_id[cid]
if cell["BlockType"] == "CELL":
grid[(cell["RowIndex"], cell["ColumnIndex"])] = (child_text(cell, by_id), cell["Confidence"])
result.append(grid)
return resultResolve relationships through the ID index, never by list position, and carry confidence and page through every transformation; a flat dictionary of strings has discarded what validation needs.
Architecture of an asynchronous pipeline
For anything beyond single-page images, the pipeline is event driven. A document lands in S3. A starter function, triggered by the S3 event or by EventBridge, calls StartDocumentAnalysis with an idempotency token. Textract processes the document and publishes a completion message, containing the job ID, status and your job tag, to an SNS topic using a role you grant it. The topic feeds an SQS queue, which buffers bursts, retries failures and parks poison messages in a dead-letter queue. A fetcher function pages through the results, normalises Blocks into records, validates them and routes each document to the database or a review queue.
Multi-step flows, such as classifying first and then analysing per document type, belong in Step Functions. Either way, set OutputConfig and KMSKeyId so raw results land encrypted in your bucket and can be re-normalised without paying again.
The pipeline in code
import hashlib, json, time, boto3
textract = boto3.client("textract")
def start(bucket, key, version_id):
# Same object version -> same token -> same JobId: S3 event redelivery cannot double-bill.
token = hashlib.sha256(f"{bucket}/{key}/{version_id}".encode()).hexdigest()[:64]
return textract.start_document_analysis(
DocumentLocation={"S3Object": {"Bucket": bucket, "Name": key, "Version": version_id}},
FeatureTypes=["TABLES", "FORMS", "QUERIES"],
QueriesConfig={"Queries": [
{"Text": "What is the invoice number?", "Alias": "INVOICE_NO"},
{"Text": "What is the total amount due?", "Alias": "TOTAL"},
]},
ClientRequestToken=token,
JobTag="invoice",
NotificationChannel={"SNSTopicArn": TOPIC_ARN, "RoleArn": PUBLISH_ROLE_ARN},
OutputConfig={"S3Bucket": RAW_BUCKET, "S3Prefix": "textract-raw/"},
KMSKeyId=KMS_KEY_ID,
)["JobId"]
def fetch_all(job_id):
blocks, token = [], None
while True:
kwargs = {"JobId": job_id, "MaxResults": 1000}
if token:
kwargs["NextToken"] = token
resp = textract.get_document_analysis(**kwargs)
status = resp["JobStatus"] # IN_PROGRESS | SUCCEEDED | FAILED | PARTIAL_SUCCESS
if status == "IN_PROGRESS":
raise RetryLater(job_id) # let SQS redeliver with a visibility timeout
if status == "FAILED":
raise JobFailed(job_id, resp.get("StatusMessage"))
blocks.extend(resp["Blocks"])
warnings = resp.get("Warnings", []) # per-page problems, also present on PARTIAL_SUCCESS
token = resp.get("NextToken")
if not token:
return status, blocks, warnings
def handler(event, context): # SQS-triggered fetcher
for record in event["Records"]:
note = json.loads(json.loads(record["body"])["Message"])
status, blocks, warnings = fetch_all(note["JobId"])
persist(normalise(blocks), status=status, warnings=warnings, job=note["JobId"])The details that matter. ClientRequestToken makes the start idempotent: the same token returns the same JobId, so a redelivered S3 event does not start and bill a second job, and reusing a token with different parameters fails with IdempotentParameterMismatchException. Deriving it from bucket, key and version ID ties the job to exactly one immutable object version. GetDocumentAnalysis returns at most 1,000 blocks per call, so a multi-page document always requires following NextToken until it is absent. JobStatus can be IN_PROGRESS, SUCCEEDED, FAILED or PARTIAL_SUCCESS; the last means some pages failed, which the Warnings list identifies by page, and it must not be treated as success.
Queries and adapters
Forms and tables give you everything Textract found; you still have to work out which key means invoice number on this supplier's layout. Queries invert that: you ask a natural-language question such as "What is the total amount due?" with an alias, and Textract returns a QUERY_RESULT for it. For variable layouts this removes most of the key-matching code, at the price of a per-query charge and the per-page query limits. Queries are English only.
When a document family is consistently hard, adapters customise the queries feature for your documents. You train an adapter on your own annotated samples (the documented dataset bounds are 5 to 2,500 training and 5 to 1,000 test documents) and pass its ID and version in AdaptersConfig. Treat an adapter like any model: version it, evaluate it on a held-out set before switching, and keep the previous version available for rollback.
Worked example: invoice extraction with validation
A finance team receives a few thousand supplier invoices a day as PDFs by email. They are saved to S3 and flow through the pipeline above with TABLES, FORMS and two queries. The expense API is the other reasonable choice and normalises common invoice fields for you; the general analysis call is used here because the team needs custom fields and the raw tables. The normaliser maps query results to INVOICE_NO and TOTAL, and takes line items from the table whose header row contains an amount column.
from decimal import Decimal, InvalidOperation
def parse_money(s):
try:
return Decimal(s.replace(",", "").replace("$", "").strip())
except InvalidOperation:
return None
def validate_invoice(rec, min_conf=90.0):
problems = []
for field in ("INVOICE_NO", "TOTAL"):
f = rec.get(field)
if f is None:
problems.append(f"{field}: missing")
elif f["confidence"] < min_conf:
problems.append(f"{field}: confidence {f['confidence']:.1f}")
total = parse_money(rec.get("TOTAL", {}).get("text", ""))
lines = [parse_money(r["amount"]) for r in rec.get("line_items", [])]
if total is not None and lines and None not in lines:
if abs(sum(lines) + rec.get("tax", Decimal(0)) - total) > Decimal("0.01"):
problems.append(f"line items do not add up to {total}")
return problems # empty -> auto-accept; otherwise route to the review queueValidation combines two signals: confidence thresholds catch fields Textract was unsure about, and arithmetic and format rules catch fields it was sure about and still got wrong. Set thresholds from data: hand-label a few hundred documents, plot accuracy against confidence, and pick the lowest threshold that meets your target. The automation rate then becomes a measured number you can improve.
Human review after A2I
Textract historically routed low-confidence results to people through Amazon Augmented AI via HumanLoopConfig. According to the Textract API reference, A2I entered maintenance mode in July 2026 and no longer accepts new customers; requests from accounts that are not existing A2I customers fail with InvalidParameterException. New systems should plan their own review step.
That is not much work: a queue table of documents, fields, confidences and problems, and a small UI that shows the page with bounding boxes from Geometry highlighted and records corrections. Those corrections are labelled data for threshold tuning and adapter training.
Failure modes
- Throttling. Transactions per second and concurrent asynchronous jobs are account quotas per Region. Starts over the job limit fail with
LimitExceededException; calls over the rate fail withProvisionedThroughputExceededExceptionorThrottlingException. Retry with exponential backoff and jitter, and cap concurrency at the starter instead of retrying a storm. - Partial success treated as success. Pages silently missing from a 200-page contract. Check
Warningsand page counts againstDocumentMetadata. - Lost notifications. Jobs finish but are never fetched. A sweeper fetches started jobs with no results after a timeout, well within the 7-day JobId validity.
- Bad input.
BadDocumentExceptionandUnsupportedDocumentExceptionfor encrypted, XFA or corrupt files;DocumentTooLargeExceptionpast the size limits. Validate type, size and page count before starting, and route rejects to a separate queue. - Silent misreads. High-confidence wrong values on poor scans; only business validation catches them.
- Permission errors.
InvalidS3ObjectExceptionusually means a role, bucket policy or KMS key denies access.
Cost, security and trade-offs
Textract charges per page at rates that depend on the API and features requested; see the current pricing page. Request only features you use, drop blank pages, and measure cost per accepted document, since review effort usually dominates.
For security, use encrypted buckets and a customer-managed output key, scope the SNS and function roles to exactly their buckets and topics, and treat extracted text as at least as sensitive as the source. Compared with a general multimodal model in Amazon Bedrock, Textract gives geometry, confidence and table structure without hosting, but fixed limits and few languages. Many teams combine them: Textract for layout and text, a language model for interpretation.
What to do next
- Collect 200 real documents, including the worst scans, and confirm they fit the format, size, resolution and language limits.
- Run them through AnalyzeDocument or AnalyzeExpense once and inspect the Blocks, not just the console view.
- Build the asynchronous pipeline with a ClientRequestToken derived from the object version, SNS to SQS with a DLQ, full NextToken pagination and handling for PARTIAL_SUCCESS.
- Hand-label the sample and set per-field confidence thresholds from measured accuracy.
- Write business validation rules and a minimal review UI; track the automation rate weekly.
- Archive raw output with OutputConfig and KMS, and add a sweeper for jobs whose notifications never arrived.