Document AI is Google Cloud's managed service for turning PDFs and scanned images into structured data: the text, its layout on the page, tables, key-value pairs and typed fields such as an invoice total or a supplier name. The API surface is small. You create a processor, send it a document, and get back a Document object. Most of the engineering effort goes into everything around that call. Which processor and which version? Online or batch? How do you trust a field that came back at 0.62 confidence? How do you keep a model upgrade from silently changing your numbers?

This article covers the processing model, the response format most teams misread, and a worked invoice-intake pipeline with real client code. Limits quoted come from the Document AI limits page at the time of writing; re-check them before you size a system.

Processors, versions and locations

A processor is a named, regional resource that wraps one extraction model. You create it once (in the console or through the API) in a location such as us or eu, and every request names it by its full resource path: projects/P/locations/L/processors/ID. The location is not cosmetic. The client must talk to the matching regional endpoint, for example eu-documentai.googleapis.com, and the documents are processed in that region. This is the first thing to settle if you have data-residency obligations.

Each processor has one or more processor versions. Google publishes pretrained versions, and for custom processors every training run you do produces a new version. A request sent to the bare processor path uses whichever version is marked as the default. A request sent to projects/P/locations/L/processors/ID/processorVersions/V uses exactly that version. That one detail decides whether your extraction is reproducible, and it gets its own section below.

Families: digitize (Enterprise Document OCR: text and layout), structure (Form Parser, Layout Parser: key-value pairs, tables, heading-aware blocks), pretrained specialised (invoice, expense, bank statement, identity documents: typed entities), and custom (Custom Extractor, Classifier, Splitter: your own schema).

The Document object and text anchors

Everything comes back as one Document message, and its central idea is simple. The whole document's recognised text is stored once, as a single string in document.text. Every other element (a page block, a paragraph, a token, a form field, an entity) does not hold its own copy of the text. It holds a text anchor: a list of segments with start_index and end_index offsets into that string. A field that spans a line break, or two columns, has several segments.

Geometry sits alongside. Each layout element has a bounding polygon in normalised page coordinates and a confidence. Pages carry their dimensions, detected languages, and for structure processors lists of form_fields and tables. Specialised and custom processors fill document.entities. Each entity has a type_, a mention_text, a confidence, page anchors pointing back at the region it was read from, an optional normalized_value (a parsed date or money amount, for example) and nested properties for repeated groups such as invoice line items.

So write one helper that turns an anchor into text, joining every segment, and keep the anchors when you store results. A value you can map back to a page region is one a reviewer can verify in seconds.

Choosing a processor

NeedStart withWhy
Searchable text from scansEnterprise Document OCRText plus layout, no interpretation; cheapest structure to reason about
Key-value pairs and tables from arbitrary formsForm ParserGeneric structure without training
Chunks for retrieval-augmented generationLayout ParserHeading-aware blocks and optional chunking
Invoices, receipts, bank statements, IDsThe matching pretrained processorA typed schema you do not have to design
Your own document class and schemaCustom ExtractorDefine fields, label samples, version the result
Mixed packets of several document typesCustom Classifier or Splitter, then routePick the right extractor per page range

For mixed inputs, classify first and route each piece to the extractor for its class. A mismatched schema produces confident nonsense rather than empty fields.

Online versus batch processing

There are two ways to call a processor. Online processing (process_document) is synchronous. You send the bytes inline and get the Document back in the response. The limits page currently lists 40 MB per file and 15 pages for most processors in this mode (the expense parser is lower). Use it for interactive flows, where a user uploads one receipt and waits.

Batch processing (batch_process_documents) is asynchronous. You name input files in Cloud Storage, either as an explicit list or as a prefix, and an output prefix. You get back a long-running operation, and results are written to the output prefix as JSON Document shards. Per-file limits are much higher: 1 GB, and page caps that depend on the processor (500 pages for Enterprise Document OCR and Layout Parser, fewer for others). The current limits page lists 5,000 files per request, though some older samples say 1,000. Large backfills and anything over the online page cap belong here.

Page caps decide the design more than file sizes. A 40-page contract cannot go online, and splitting it yourself breaks entities that cross the cut.

An online call and the anchor helper

Here is a complete online call with the google-cloud-documentai Python client, plus the anchor helper every consumer needs. The endpoint must match the processor's location.

from google.api_core.client_options import ClientOptions
from google.cloud import documentai_v1 as documentai

PROJECT, LOCATION = "my-project", "eu"
PROCESSOR, VERSION = "YOUR_PROCESSOR_ID", "YOUR_PROCESSOR_VERSION_ID"

client = documentai.DocumentProcessorServiceClient(
    client_options=ClientOptions(api_endpoint=f"{LOCATION}-documentai.googleapis.com"))

# Pin the version: the bare processor path follows whatever default is set later.
name = client.processor_version_path(PROJECT, LOCATION, PROCESSOR, VERSION)

def anchor_text(doc, anchor):
    """Join every segment; start_index is omitted from the message when it is 0."""
    return "".join(doc.text[int(s.start_index):int(s.end_index)] for s in anchor.text_segments)

with open("invoice.pdf", "rb") as f:
    raw = documentai.RawDocument(content=f.read(), mime_type="application/pdf")

result = client.process_document(request=documentai.ProcessRequest(name=name, raw_document=raw))
doc = result.document
for e in doc.entities:
    page = e.page_anchor.page_refs[0].page if e.page_anchor.page_refs else None
    print(e.type_, repr(anchor_text(doc, e.text_anchor)), round(e.confidence, 2), "page", page)

Fill in the version id by listing the real ones with client.list_processor_versions(parent=processor_path). Entity type names such as total_amount come from the processor schema; do not hard-code names you have not seen in a real response.

Batch processing and reading the shards

The batch path has three parts: submit, wait, then read every shard. One input file can produce several output JSON files, so never assume one-to-one.

import re
from google.cloud import storage

def submit_batch(input_prefix, output_prefix):
    request = documentai.BatchProcessRequest(
        name=name,
        input_documents=documentai.BatchDocumentsInputConfig(
            gcs_prefix=documentai.GcsPrefix(gcs_uri_prefix=input_prefix)),
        document_output_config=documentai.DocumentOutputConfig(
            gcs_output_config=documentai.DocumentOutputConfig.GcsOutputConfig(gcs_uri=output_prefix)),
    )
    return client.batch_process_documents(request=request)   # long-running operation

def read_results(operation, timeout=1800):
    operation.result(timeout=timeout)                 # raises if the whole operation failed
    meta = documentai.BatchProcessMetadata(operation.metadata)
    gcs = storage.Client()
    for status in meta.individual_process_statuses:
        if status.status.code != 0:                   # per-file failure: record it, keep going
            yield status.input_gcs_source, None, status.status.message
            continue
        bucket, prefix = re.match(r"gs://(.*?)/(.*)", status.output_gcs_destination).groups()
        for blob in gcs.list_blobs(bucket, prefix=prefix):
            if blob.name.endswith(".json"):
                shard = documentai.Document.from_json(blob.download_as_bytes(),
                                                      ignore_unknown_fields=True)
                yield status.input_gcs_source, shard, None

A batch operation can succeed overall while individual files fail, so the per-file status loop is not optional. In a real service, store the operation name and poll from a separate job rather than blocking on operation.result(), so a restart does not lose track of running work.

Worked example: invoice intake

Take a finance team that receives about 3,000 supplier invoices a day as PDFs by email. 92 percent are one to three pages, and a long tail of statements runs past 40 pages. The goal is supplier, invoice number, date, currency, total and line items in BigQuery, with no more than 5 percent of invoices needing a person.

The pipeline in the diagram does the following. An email gateway drops attachments into an inbox bucket. An Eventarc trigger on object finalisation starts a Cloud Run job. To avoid one operation per file, the job collects files for a few minutes and submits one batch request per group to a pinned invoice-parser version, using a per-group output prefix. When the operation completes, a validator reads each shard and applies checks the model cannot make for you. Line-item amounts must sum to the subtotal within a rounding tolerance. Tax plus subtotal must equal the total. The invoice number must match the supplier's known pattern. The date must be no later than today. The supplier must resolve to a row in the vendor master.

Any rule failure goes to review, as does any required field below a per-field confidence threshold set from 500 hand-checked invoices, not guessed. The review screen highlights the anchored region, and corrections become labelled examples for the next evaluation.

Invoice intake: files land in a bucket, batch jobs parse them, validation decides what a person checksInbox bucketPDF, TIFF, PNGEventarc triggerobject finalizedCloud Run jobgroups files, submitsDocument AIpinned processor versionbatch requestOutput bucketDocument JSON shardswritesValidatoranchors, sums, rulesLRO doneBigQueryaccepted fieldsReview queuelow confidence or rule failpassrouteLabelled samplesfrom reviewer fixescorrectionsEvery stored value keeps its page and text offsets, so a reviewer can see exactly where it came from.
Event-driven invoice intake: batch processing against a pinned version, deterministic validation, and a review loop that also produces labelled data.

Pinning and evaluating processor versions

Pretrained processors get new versions, and the default can move. If your code calls the bare processor path, a default change can alter field boundaries, date normalisation or line-item grouping overnight with no deploy on your side. Treat the processor version like any other dependency:

  1. Pin the full version path in configuration, never in a code literal that nobody reviews.
  2. Keep a frozen evaluation set of real documents with ground-truth fields, stratified by supplier and layout.
  3. When a new version appears, run the set against both versions and compare per-field accuracy, not one aggregate score. A version that is better overall can be worse on the one field your ledger depends on.
  4. Promote by changing configuration, then watch review-queue rate and rule-failure rate for a week.

Custom processors need the same discipline: every training run is a new version.

Layout Parser and chunking for retrieval

For retrieval-augmented generation, the Layout Parser is the processor to know. Plain OCR loses the fact that a sentence sits under the heading "Termination" in section 9. The layout parser keeps a block hierarchy (headings, paragraphs, lists, tables), and it can split the document into chunks. You configure chunking through process options:

opts = documentai.ProcessOptions(
    layout_config=documentai.ProcessOptions.LayoutConfig(
        chunking_config=documentai.ProcessOptions.LayoutConfig.ChunkingConfig(
            chunk_size=500,                 # tokens per chunk
            include_ancestor_headings=True, # prefix each chunk with its heading path
        )))
request = documentai.ProcessRequest(name=layout_parser_name, raw_document=raw, process_options=opts)

Including ancestor headings is what makes a chunk retrievable on its own. "The notice period is 30 days" is ambiguous, but the same text under "Master Agreement, Termination" is not. Inspect a real response to find the chunk fields before you write the indexer, and record the source page range with each chunk so answers can cite it. Check that table chunks keep their header row.

Failure modes

  • Wrong regional endpoint. An EU processor called through the US endpoint fails with an error that looks like a permissions problem. Derive both from one location setting.
  • First-segment slicing. Multi-line addresses and wrapped descriptions come back truncated. Join all text segments.
  • Unpinned version drift. Output changes with no deploy. Pin, and alert on a jump in review rate.
  • Page-cap rejections. Online calls on long files fail. Route on page count before calling, and send long files to batch.
  • Partial batch failure treated as success. Reconcile input count against per-file statuses.
  • Confidence treated as probability. A 0.9 field is not right 90 percent of the time. Calibrate thresholds per field on labelled data.
  • Quota bursts. A backfill throttles live traffic. Rate-limit backfills and retry with backoff.
  • Poor scans. Track OCR confidence per page and bounce unreadable inputs early.

Security and operations

Grant the calling service account the Document AI API user role (roles/documentai.apiUser), and give it object read and write only on the input and output buckets. For sensitive documents, put the project inside a VPC Service Controls perimeter, so processed JSON cannot be written to a bucket outside it, and use customer-managed encryption keys where your policy requires them. Choose the processor location to match residency requirements. Set lifecycle rules on the output bucket, because those shards contain the full extracted text and quickly become the largest copy of sensitive data you hold.

Monitor per processor version: pages, per-file error rate, latency, review-queue rate and rule-failure rate by rule; the last two are your real quality metrics. Pricing is per page and differs by processor, so check the current pricing page rather than a figure in a design document.

Trade-offs

ChoiceGainsCosts
Pretrained processorNo labelling, typed schema, fast startSchema fixed by Google; versions move
Custom ExtractorYour schema, versioned training runsLabelled samples and an evaluation set to maintain
General multimodal model with a promptFlexible schema, handles odd layoutsNo anchors to the source region by default; harder to pin and audit
Online processingSimple, immediatePage and size caps; synchronous quota pressure
Batch processingLarge files, high volumeAsynchronous plumbing; shard handling; per-file status

If a wrong number has a cost, you want values anchored to a page region, a pinned model version, and deterministic validation around them. Document AI gives you the first two; the third is your job.

What to do next

  1. Measure the page-count and file-size distribution of a week of real inputs, and decide online versus batch from it.
  2. Create the processor in the region your data must stay in, and derive the endpoint from the same setting.
  3. List the processor versions, pin one by full path in configuration, and record why it was chosen.
  4. Write the anchor-to-text helper and a small script that prints entities with page numbers for ten real files.
  5. Hand-label 200 to 500 documents and set per-field confidence thresholds from them.
  6. Add deterministic checks (sums, dates, master-data lookups) and route failures to a review screen that shows the anchored region.
  7. Lock down IAM, the bucket lifecycle and the service perimeter before the first production document arrives.
  8. Re-run the frozen evaluation set whenever a new processor version is published.

Related reading on this site: Cloud Vision AI for image-level detection, Eventarc for the trigger, Workflows for orchestrating long-running operations, VPC Service Controls for the data perimeter, and the Gemini API for the multimodal alternative.

Key takeaway: Document AI returns one text string plus anchored structure, so resolve every segment and keep the anchors. Pick the processor by document class, pin its version by full path, and choose online or batch from your real page counts. Wrap extraction in deterministic checks and calibrated confidence thresholds, route the rest to review, and re-evaluate on a frozen set before any version change.