All posts
ocrself-hosted-aicompliancedocuments

OCR and Scanned PDFs: Map the Whole Document Path

Before calling a document workflow private, I check where OCR, embeddings, logs and backups actually run. A local model alone does not settle it.

Stéphane Lepain··Updated ·4 min read

Title and description revised 24 September 2026. The body and its sources were not rechecked in that revision.

Body revised 13 September 2026: this article describes a general design and test method. No private operating result, delivery time or performance guarantee is claimed.

Self-hosting a language model does not establish where documents are processed. OCR, embeddings, logs, storage and backups can use different services. I map the complete path before describing a document workflow as private.

Map the complete document path

I check which parts of the actual workflow are local and which use external providers.

For each stage, record what information it receives, where it is processed, what persists and who can access it. Include extracted text and derived indexes, not just the original PDF. An approved external provider may be appropriate; it needs to be an explicit decision. See data-processing responsibilities.

The non-negotiable rule

The processing must follow the agreed data boundary and retention requirements. There is no universal requirement that all document bytes exist only in RAM, and choosing RAM does not prove that no copy persists elsewhere.

Temporary files, swap, crash dumps, logs, queues and backups may all matter. Verify the actual configuration and failure behaviour. A successful deletion response or empty query is evidence about that interface, not proof that every copy is gone.

Stage 1: OCR — layout engine plus vision escalation

I test representative scans, text pages, tables and diagrams before choosing an extraction method. Compare the extracted text and reading order with the visible page, retaining enough source references for a person to check the result.

Different extraction tools can fail differently. A long output is not proof that the page was read correctly, and model confidence is not a substitute for comparison with the source. Mark missing or uncertain content for review.

For available capabilities and licences, consult the current Tesseract documentation and the documentation of any other proposed engine. A particular model, GPU or language configuration needs its own acceptance test.

Stage 2: Embedding — local model, project-scoped index

An external embedding service receives the text sent to it. Local embedding is an option when required, but it still needs access controls and a defined retention policy.

Check isolation using users who should and should not have access to each document. A project label or separate collection is a design component, not proof that all access paths are protected. Include retrieval, exports and administrative interfaces in the agreed checks.

Treat indexes and chunk metadata as derived information that also needs protection. Do not assume that deleting the source PDF removes its searchable content.

Stage 3: The cleanup contract

Document what must be removed, when, from which active stores and under what backup policy. Define what evidence the deletion procedure can produce and what it cannot establish.

Test the procedure on synthetic documents, including interrupted jobs and failed cleanup. Record outstanding copies or retention constraints rather than promising immediate universal erasure.

What this buys you, in board-meeting language

The aim is a workflow whose processing locations, access rules, review responsibilities and running costs can be explained and checked. It is not an automatic compliance certification or a guarantee of perfect extraction.

Compare the effort of checking the result with the current manual task. Include extraction, embedding, hosting, human review, correction, maintenance and support.

The honest caveats

Speed, memory needs and extraction quality depend on the inputs, tool versions and operating setup. I do not quote a universal local-versus-cloud speed ratio or a fixed delivery time before scoping the work.

A pilot should include difficult pages, missing information and cases that need a person. Agree the document types and acceptance criteria before building. Describe one recurring document task without sending confidential files in the first message.

Questions people ask about sovereign document pipelines

Why does OCR matter if the model is local? It is a separate processing step with its own data destination and retention behaviour.

Must everything run locally? That depends on the agreed requirements. Every external route needs permission and appropriate provider review.

Does an empty index prove deletion? It verifies the result of that query. Other stores, replicas, backups and derived material need their own checks.

How do I know whether the extraction is good enough? Compare representative outputs with their source pages and use agreed acceptance criteria, including unresolved cases.