Skip to content
Apsan Works

Automating Document Processing With AI: A Practical Guide

Invoices, contracts, applications, claims. The architecture that works, the failure modes nobody warns you about, and how to know when it is ready.

6 min readUpdated 22 September 2026

Document processing is the most common automation request we get, and the one where expectations and reality diverge most sharply. The demo, upload a clean PDF and watch structured data appear, is genuinely easy now. The production system is not, and the gap sits almost entirely in the parts nobody demos.

The pipeline that works

Six stages, each one a durable checkpoint, so a failure at stage five doesn’t discard the work done in stages one through four. The reasoning behind that design, and when it is worth applying, is covered in durable pipeline architecture.

Ingestion comes first. Documents arrive from email, a watched folder, an upload, or an API, and this stage does exactly one job: get the artefact into durable storage, record where it came from, and emit an event. Never parse here. When something goes wrong downstream, you want the original file exactly as received.

Normalisation turns everything into a consistent internal representation. Native PDFs get text extracted directly. Scans go through OCR. Images get deskewed and contrast-adjusted. Office formats get converted. Whatever came in, the output of this stage looks the same.

Classification asks the cheap question first: what kind of document is this? It’s fast and enormously valuable, because it lets everything downstream specialise. A model with a short prompt handles this fine, and a small, cheap model handles it well enough.

Extraction is the actual work: pulling the fields you need into a defined schema. Quality here depends far more on how you frame the task than on which model you use. Ask for a strict schema, keep the document’s layout information where you can, and require a null rather than a guess when a field is genuinely absent. The same schema-first habit carries to long solicitations, as in RFP response automation, where extraction has to feed a bid decision.

Validation checks the extracted data against rules that have nothing to do with the model. Do the line items sum to the total? Is the date within a plausible range? Does the reference number match the expected format? Does this supplier even exist? Most extraction errors get caught here, cheaply and deterministically.

Routing decides what happens next. Confident and valid goes straight through. Anything else lands in a review queue, with the specific problem flagged and the source document displayed next to the extracted values.

The failure modes nobody mentions

Some documents aren’t really documents. A “PDF invoice” that’s a photo of a screen. A scan at 150 DPI where the decimal points are ambiguous. A fax-to-PDF that’s been through three generations of compression. These are common in exactly the industries most eager to automate, and no amount of prompting fixes an unreadable input. Detect it and route it rather than attempting heroics.

Multi-document files cause quiet damage. One PDF containing an invoice, two delivery notes, and a signed page from something unrelated. Assume one file equals one document and this corrupts your data silently, so splitting deserves to be its own classification step.

Tables that span pages are a genuinely hard layout problem, with line items broken across a page boundary and the header sometimes repeated, sometimes not. This is where most extraction accuracy gets lost on complex documents.

The most dangerous failure is values that look right and aren’t, a confidently extracted total pulled from the wrong column. It passes every check you didn’t think to write, which is exactly why validation rules should encode real business logic (sums, cross-references, plausible ranges) rather than just confirming that a number was found somewhere.

Format drift is the quiet one. A supplier changes their invoice template, and your accuracy drops for that supplier alone, with no error raised anywhere. Per-source accuracy monitoring catches this. Aggregate monitoring does not.

Confidence and the review queue

The single most important design decision is what happens when the system is unsure, and the answer is never to guess.

Confidence should come from several signals combined, not from the model’s self-report alone. Did validation pass? Are the extracted values internally consistent? Does this document match a known template? How does the extraction compare to previous documents from the same source? Together, these are far better calibrated than any one on its own.

Route on that combined signal. High confidence proceeds automatically. Everything else lands in a review queue, and the quality of that queue decides whether the whole project succeeds. A good one shows the document and the extracted fields side by side, highlights the specific field that failed, and lets someone fix it in a keystroke or two rather than a minute. A bad one is a spreadsheet of rejected items nobody opens.

The corrections from that queue are also your best source of training and evaluation data, so capture them.

The same routing logic, confident cases through, uncertain ones flagged with context, shows up again in building a monitoring agent that does not cry wolf, applied to alerts instead of extracted fields.

What accuracy to expect

Be sceptical of any number quoted without a document set attached to it. Accuracy on clean, consistent, native-PDF documents from a handful of known sources runs very high. Accuracy on mixed-quality scans from hundreds of sources runs meaningfully lower. Both get marketed as “document AI.”

The honest framing isn’t a percentage. It’s what proportion of documents can go through without a human touching them, and how long the remainder takes to review. A system that clears seventy percent automatically and makes the other thirty percent take fifteen seconds each is transformative. A system that clears ninety-five percent but makes the remaining five percent take ten minutes of investigation each may be worse than what you had before.

Rolling it out without risking anything

Run in shadow first. The pipeline processes real documents alongside the existing manual process, and nothing it produces gets used yet. Compare outputs. Every disagreement is either a bug or a place where your own definition of correct was less settled than you thought, and you want to find both before cutover, not after.

Then go segment by segment. Start with your highest-volume, most consistent document source, since that’s where accuracy will be best and value highest, and widen only once the numbers hold.

Keep the manual path available throughout. The confidence to switch off a working process comes from watching the new one match it for a few weeks, not from a go-live date written into a plan.

This is a large part of what we build. See our approach to automation, or read about when custom beats no-code.

Tell us what is slowing you down

A short conversation is usually enough to tell whether this is a build, an automation, or something you should not do at all. We will tell you which.