Every company processes documents that someone retypes into a system: supplier invoices, delivery notes, acts of completed work, contracts, bank statements, customs declarations. Traditional OCR with templates worked for a few high-volume formats and broke whenever a supplier changed their layout. Vision-capable language models change the economics: one pipeline can read hundreds of layouts, in Ukrainian and English, from clean PDFs and phone photos. This guide describes how we build such pipelines for production — and why the extraction model is the easy part.
The pipeline at a glance
Inbox / upload / scanner
│
▼
1. Intake → deduplicate, detect type, split multi-document PDFs
2. Preprocess → text layer or page images, rotate, enhance scans
3. Extract → LLM with schema (structured output)
4. Validate → business rules, cross-checks against ERP data
5. Decide → auto-post / human review / reject
6. Integrate → create draft records in ERP, attach original
7. Learn → corrections feed the eval set and prompts
Steps 4–7 decide whether the system saves time or creates a new category of errors.
Step 1–2: Intake and preprocessing
Classify first. Before extraction, determine the document type (invoice, credit note, act, contract) with a cheap classification call. Different types need different schemas and rules.
Use the text layer when it exists. Digitally generated PDFs contain text; extracting it (with layout) is cheaper and more accurate than OCR. Libraries such as Docling convert PDFs, including tables, into structured Markdown or JSON that models read well.
Send images for scans and photos. Modern models accept images and PDFs directly and read both text and visual layout (Anthropic PDF support; vision). For poor-quality scans, basic enhancement — deskewing, contrast, sufficient resolution — measurably improves results. Very long documents should be split by page ranges.
Handle multi-document files. A single scanned PDF often contains several invoices. Detect boundaries before extraction, or ask the model to return an array of documents with page ranges.
Step 3: Schema-based extraction
Define the target data precisely with a schema and use structured output so responses always parse:
from datetime import date
from decimal import Decimal
from pydantic import BaseModel, Field
class Line(BaseModel):
description: str
quantity: Decimal
unit: str | None = Field(None, description="e.g. шт, кг, год, послуга")
unit_price: Decimal = Field(description="Price per unit excluding VAT")
vat_rate: Decimal | None = Field(None, description="VAT rate in percent, e.g. 20")
amount: Decimal = Field(description="Line total excluding VAT")
class SupplierInvoice(BaseModel):
supplier_name: str
supplier_edrpou: str | None = Field(None, description="8-digit EDRPOU or 10-digit RNOKPP, digits only")
invoice_number: str
invoice_date: date
currency: str = Field(description="ISO code, e.g. UAH, EUR")
lines: list[Line]
total_without_vat: Decimal
vat_amount: Decimal | None
total_with_vat: Decimal
iban: str | None = Field(None, description="UA + 27 digits if present")
evidence: dict[str, str] = Field(
default_factory=dict,
description="For key fields, the exact text from the document they were taken from",
)
Design principles from structured outputs from LLMs apply: nullable fields for information that may be absent, clear formats in descriptions, and an evidence field that quotes the source text. Evidence makes human review much faster and reduces invented values, because the model must point to where a value came from.
Prompt guidance specific to documents:
- Explain local conventions: "Ukrainian invoices often show amounts with a comma as decimal separator and a space as thousands separator."
- Tell the model what to do with handwriting, stamps and crossed-out values.
- Ask it to return
nullrather than guess, and to never compute totals that are not printed — computation happens in code.
Step 4: Validation — where reliability comes from
Model output is a proposal. Code decides whether it is plausible:
| Check | Example |
|---|---|
| Arithmetic | Sum of line amounts equals total_without_vat; VAT equals rate × base within rounding |
| Format | EDRPOU length and checksum; IBAN checksum; dates within a plausible range |
| Master data | Supplier exists in ERP by EDRPOU; IBAN matches the supplier's known accounts |
| Business rules | Invoice references an open purchase order; amounts within PO tolerance |
| Duplicates | Same supplier + number + date already processed |
The IBAN check deserves special attention: a changed bank account on an otherwise normal invoice is a classic fraud pattern. Never auto-update bank details from a document.
Step 5: Decide — automation with a human in the loop
Not every document should be auto-posted. Route based on validation and risk:
- Auto-create a draft when all checks pass and the supplier is known.
- Human review when any check fails, a field is null, the supplier is new, or the amount exceeds a threshold. Show the document side by side with extracted fields and highlighted evidence.
- Reject or escalate when the document is unreadable or suspicious.
Model "confidence" scores are not well calibrated; base routing on deterministic validation rather than on asking the model how sure it is. Track the share of documents processed without corrections — that is your real automation rate. For complex exceptions such as unknown suppliers or PO mismatches, an agent with read-only tools can prepare a resolution for the reviewer; see workflows vs agents.
Step 6: ERP integration
Create draft records, never posted ones, and attach the original file and the extraction result for audit. For Odoo, the external API lets you create vendor bills with lines, link partners by tax ID and attach PDFs; our integration patterns are in Odoo AI integration with the JSON-2 API, and local payment and logistics integrations in Odoo with Nova Poshta and Monobank. Odoo Enterprise also has its own document digitisation features; compare them with a custom pipeline before building — see Odoo Community vs Enterprise.
Make the integration idempotent: a document ID or content hash prevents duplicate bills when a job is retried; see LLM API reliability.
Step 7: Measure and improve
Build a labelled set of 200–500 real documents across suppliers and quality levels, and measure:
- Field accuracy per field (exact match after normalisation).
- Document-level accuracy: share of documents with all critical fields correct.
- Null recall: when a value is absent, does the model return null rather than invent one?
- Straight-through rate: share of documents that pass validation without human edits.
- Time per document for reviewers, before and after.
Every human correction is a labelled example. Feed a sample into the eval set weekly and re-run it on every prompt or model change; see LLM evals.
Costs and throughput
- A typical one- or two-page invoice costs a fraction of a cent to a few cents to extract, depending on the model and whether you send text or images.
- Use a cheaper model for classification and clean digital PDFs; a stronger model for scans, handwriting and complex tables.
- Process in a queue with concurrency limits; use batch APIs for non-urgent backlogs — see LLM cost optimization.
- Compare with the baseline: the fully loaded cost of manual entry is usually dollars per document, not cents.
Privacy and compliance
Invoices and contracts contain personal data (sole proprietors' names and tax numbers, signatures) and commercial secrets. Use providers with appropriate data processing terms and zero or limited retention, keep originals in your own storage, restrict access to the review UI, and set retention periods. For strict requirements, extraction can run on self-hosted vision models. Details are in privacy for LLM apps. In the EU, check whether any downstream decision (for example, credit checks) falls under high-risk categories of the AI Act.
A rollout plan
- Weeks 1–2: collect 300 representative documents, label 100, define schemas and validation rules.
- Weeks 3–4: build the pipeline with drafts only and a review UI; measure accuracy on the labelled set.
- Month 2: pilot with one team; all documents reviewed; collect corrections.
- Month 3: enable straight-through processing for known suppliers with all checks passing; keep sampling audits.
FAQ
Do we still need OCR? For digital PDFs, use the text layer. For scans, vision models read images directly; a dedicated OCR step helps mainly for very poor scans or when you need searchable text archives.
How accurate is it? On clean documents and well-designed schemas, field accuracy is typically very high; scans, handwriting and unusual layouts lower it. Measure on your own documents — there is no universal number.
Can it handle Ukrainian and mixed-language documents? Yes, strong multilingual models handle Ukrainian, English and Russian documents and mixed layouts. Include examples of each in your eval set.
What about contracts and long documents? Extract specific fields (parties, term, amounts, termination clauses) with evidence quotes, and process by section for long documents. Treat legal interpretation as decision support, not automation.
Sources
- Anthropic. PDF support and vision.
- OpenAI. Structured outputs.
- Docling — document conversion for AI.
- Kim et al. (2022). OCR-free Document Understanding Transformer (Donut).
- Huang et al. (2022). LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking.
- Odoo. External API documentation.