Applied document understanding

Create a Document AI Processor

Example: turn invoice and receipt documents into searchable text and reviewable business fields. OCR supplies text and layout; specialized processors can also return typed entities.

Learning only. Documents, geometry and extracted fields below are hand-authored fixtures, not API results. No files are uploaded or processed.

FIXED DOCUMENT / LIVE ANCHOR SELECTION

Document → text → structure

FICTIONAL SAMPLE / NOT A SCANNORTHSTAR SUPPLIESInvoice INV-1042Date: 2026-10-01Customer: Example StudioNotebooks x 10 USD 50.00Pens x 20 USD 20.00Subtotal: USD 70.00Tax: USD 7.00Total: USD 77.00PAGE 1 / SYNTHETIC GEOMETRY
Page 1 · 612 × 792 pt · normalized layout coordinates

TEXT ANCHOR / GLOBAL OFFSETS

NORTHSTAR SUPPLIES

[0, 18) · end is exclusive

NORTHSTAR SUPPLIES
Invoice INV-1042
Date: 2026-10-01
Customer: Example Studio
Notebooks x 10    USD 50.00
Pens x 20         USD 20.00
Subtotal:        USD 70.00
Tax:              USD 7.00
Total:           USD 77.00

Dimensions carry their own returned unit; do not label every page as pixels. The toy offsets use ASCII text; the Python code follows the API’s full-document text anchors.

Online or batch processing

process_document

Input

One supported document as raw bytes or a supported Cloud Storage input.

Result

Waits for the processed Document response. Suitable for bounded, on-demand work.

Limits & safety

Page, size and timeout limits depend on the processor, version and options. Check current limits; there is no universal 15-page rule.

Google Cloud Document AI setup

Choose a processor

Select a project and supported processor location. Use Enterprise Document OCR for text and layout; use a suitable Invoice or Expense processor for structured business fields. Create or select a processor in the Document AI console and copy its actual resource ID.

Your terminal
gcloud services enable documentai.googleapis.com --project=YOUR_PROJECT_ID

Run only in your own project after reviewing billing, IAM and pricing. Creating processors and processing files may incur charges; budget alerts are not spending caps.

Python · real online processing

Run only in your own authenticated environment. This synchronous example sends document bytes, reads text anchors and preserves page dimension units; it is not executed here.

Python & C++ examples

Python dependencies: google-cloud-documentai · ADC · billable API when run externally
import argparse
import json
from pathlib import Path
from google.api_core.client_options import ClientOptions
from google.api_core.exceptions import GoogleAPICallError
from google.cloud import documentai


def anchor_text(text, anchor):
    return "".join(text[int(segment.start_index):int(segment.end_index)]
                   for segment in anchor.text_segments)


def process_document(path, project, location, processor, mime_type):
    content = path.read_bytes()
    if not content:
        raise ValueError("Document is empty")
    # ADC supplies credentials; match your processor's regional endpoint.
    client = documentai.DocumentProcessorServiceClient(
        client_options=ClientOptions(
            api_endpoint=f"{location}-documentai.googleapis.com"
        )
    )
    try:
        name = client.processor_path(project, location, processor)
        request = documentai.ProcessRequest(
            name=name,
            raw_document=documentai.RawDocument(
                content=content, mime_type=mime_type
            ),
        )
        document = client.process_document(request=request, timeout=120).document
        return {
            "text": document.text,
            "pages": [{
                "page_number": page.page_number,
                "width": page.dimension.width,
                "height": page.dimension.height,
                "unit": page.dimension.unit,
                "text": anchor_text(document.text, page.layout.text_anchor),
            } for page in document.pages],
            "entities": [{
                "type": entity.type_,
                "mention_text": entity.mention_text,
                "anchored_text": anchor_text(document.text, entity.text_anchor),
                "confidence": entity.confidence,
            } for entity in document.entities],
        }
    finally:
        client.transport.close()


if __name__ == "__main__":
    parser = argparse.ArgumentParser(description="Document AI online processing")
    parser.add_argument("file", type=Path)
    parser.add_argument("--project", required=True)
    parser.add_argument("--location", required=True)
    parser.add_argument("--processor", required=True)
    parser.add_argument("--mime", required=True,
                        choices=["application/pdf", "image/jpeg", "image/png", "image/tiff"])
    args = parser.parse_args()
    try:
        result = process_document(args.file, args.project, args.location,
                                  args.processor, args.mime)
        # Contains document content: print only in a private environment.
        print(json.dumps(result, indent=2, ensure_ascii=False))
    except (OSError, ValueError, GoogleAPICallError) as error:
        parser.exit(1, f"Document processing failed: {error}\n")

Before using extracted documents

01Use consented documents, restrict processor and storage access, and check regional availability and data handling requirements.

02Evaluate field accuracy and missing values against representative ground truth; confidence is not a guarantee of correctness.

03Validate MIME types, content and current processor limits. Handle errors and operation state without duplicate chargeable submissions.

04Keep raw document content out of public logs. Require human review before invoices trigger payments or other consequential actions.