→How It's Built

Check the evidence before trusting the result.

Start with what I've used in my day-to-day work for two years, preparing reviews for specialists to sign off. Then check the fictional source example below: it shows four specific behaviours, not general AI reliability or customer savings. The technical sections explain choices to assess for a build; they are not evidence that every proposed workflow is already proven.

00Production evidence

One person prepares what a team of specialists signs off.

I prepare alone the review that a team of 3 to 5 specialists signs off: security architects for each domain, a programme manager and a product owner.

For two years, I've used a system I built in my day-to-day work. It reads long documents, checks them against the requirements that apply, and writes up what's missing with the page cited. A second AI reviews the first one's findings independently before I see them. In a 50-case test, the two AIs agreed on all 50 verdicts, and the second never passed a wrong 'compliant'. Fifteen automated checks run on every review before I do. Measured in September 2026: 2,773 pages read, across 38 document versions.

2 years

in my day-to-day work, one person

2,773 pages

read across 38 versions, Sep 2026

50/50

verdicts matched in a 50-case test

15 checks

run on every review before I see it

What this does and doesn't show: these are measured figures from one system I built and use myself, not a customer case study with a name attached. The organisation whose work I do is never identified publicly, by agreement; I'm the only person who operates the system. The diagram-recovery and flag-rate figures describe the material they were measured on, not every document ever processed, and the flag rate is not an error-detection recall rate or an accuracy percentage. There is no measured uptime, user count, or cost-saving figure to publish alongside these; none exists in a form I can state plainly, so none is claimed here.

One of the fifteen checks exists because of a real failure, not a hypothetical one: when an AI validator invented a rule that wasn't there, and how I caught it before it reached a decision. That's what "I review and decide" means in practice.

01Synthetic example

A citation you can check yourself.

This is a separate capability demonstration, not the production system above. I wrote a fictional two-page workshop handbook for this example. It contains no customer material. Page 1 says: “Refunds are issued within 10 working days.”

A Python script checks a supplied quotation against the supplied page. In the four included cases, the real quotation is found, an invented two-day quotation is rejected, the correct quotation on the wrong page is rejected, and an unanswered accessibility question is left for a person.

present-quote: quote-present
invented-quote: quote-not-found
wrong-page: quote-not-found
missing-information: needs-human-answer
4/4 fixture checks passed

Reproduced on 9 September 2026 with the published files. Download the handbook and cases and Python checker into one folder, then run python3 source-check.py. It uses the standard library and makes no network calls.

What this proves: these four supplied cases behave as shown. It does not generate an AI answer, prove that a conclusion follows from a quote, validate OCR, or establish production accuracy. In a build, those need separate tests with your agreed acceptance criteria.

I check three recorded answers from fictional documents in this video, using my AI likeness and voice; loading the YouTube player connects to Google.

For the broader workflow, see how I scope assistance and automation.

02Architecture

One place to control model access.

I separate the workspace from its models so a model can change without rebuilding how your team works. I agree the allowed routes with you and check them for each build, including any external API.

  • · The stack I use. Open-source LibreChat for the workspace, a LiteLLM gateway for model choice, and Ollama for local open models.
  • · Your choices. External providers can be connected with your own API keys if you approve them. Routes, access and usage need checking for each build.
  • · Your handover. I write down the configuration and walk your team through it.
Chat workspaceyour team's daily interfaceDocument pipelineOCR, search, workflowsYour applicationsone OpenAI-compatible APIONE ROUTERrouting policy · per-call logs · cost tracking · budgetsAGREED ROUTES — TEST BEFORE USESENSITIVE ROUTESALLOWED ROUTESLocal models — your hardwarework agreed for local processingverify the complete data boundaryFrontier providersapproved information and providersverify provider terms and usage logs
Illustrative design. I document the permitted routes and test the configuration and logs before relying on this boundary.

Model selection and data access are separate decisions. A difficult question must not be sent outside an agreed boundary just because a larger model might answer it better.

03Documents

Check what was read before trusting the answer.

A successful upload does not mean a document was read correctly. Scans, tables and diagrams need different handling. I start with representative documents and check the extracted result against the pages a person can see.

Before calling a document workflow private, I map the whole document path, including extraction, indexes, logs and backups.

  • · Choose extraction by page type. Use the text layer where it is available, OCR for scans, and appropriate handling for diagrams.
  • · Show gaps. Flag unreadable or incomplete pages for review. Missing extracted text is not evidence that the source contains nothing.
  • · Keep the source traceable. Preserve document versions and page references so a reviewer can check a finding.

I agree with you what acceptable extraction looks like before relying on it downstream. Documents that fail those checks stay marked for review.

04Review

Give the reviewer something they can check.

For document-review work, I agree how a person checks the draft against its sources and which checks can be performed by code. The checks need to fit the task and be tested on agreed examples.

  • · Check citations against the source. A plausible page number is not enough; the cited passage must be there.
  • · Keep disagreements visible. Conflicting findings go to your reviewer. I agree how unresolved cases reach the person responsible.
  • · Use your rules. Checks refer to an agreed source. A model cannot invent a requirement to justify its answer.
  • · Render from structured findings. Use a consistent format so a person can compare and review the output.

For repeatable deliverables, I keep findings separate from the document that presents them. A person still checks the reasoning and the exported result.

A correct citation does not make the conclusion correct. Your reviewers still make the judgment. I test the workflow against agreed examples, including cases it should reject or leave unresolved.

05Operations

Your team should be able to take over.

  • · Backups and recovery. I agree with you the recovery and data-loss targets for your environment, then test the restore procedure.
  • · Monitoring. Your team needs to see failed services, storage pressure and API spend, with a clear response for each alert.
  • · Controlled changes. Configuration is versioned. Updates have checks and a way back if they fail.
  • · A practical handover. Your team follows the runbook with me. Gaps found during that walkthrough are part of the work to finish.

Recovery times depend on the hardware, data and backup arrangement. I agree them for your build rather than quote a result from another system. Ongoing support can be scoped separately if you want it.

06Limits

What I agree with you before a build.

  • · Whether private AI is needed. If an existing service covers the job and meets your requirements, I will say so.
  • · What the hardware can handle. Model size, document length and simultaneous users affect memory needs and response time.
  • · What may leave the environment. External APIs receive the data sent to them. Provider terms and routing rules both need checking.
  • · Who owns the decisions. I agree with you who approves data access, reviews outputs and operates the system after handover.

Tell me what your team needs to do.

A free forty-five-minute scoping call ends with a short written note: what you could run, what it would take, and whether it is worth pursuing. Any build gets a written scope and an agreed price first.