All posts
ai-pilotai-roiprivate-aicost-controldocument-workflows

AI Pilot Cost: Measure Cost per Accepted Result

Define an accepted result, count failed attempts and human review, and compare the full cost of one recurring task before committing to a build.

If checking the answer takes longer than doing the work, I want that on the spreadsheet.

Stéphane Lepain··Updated ·6 min read

To evaluate an AI pilot, divide the full cost of the agreed test by the number of results that pass its acceptance checks. I use this approach when scoping AI assistants and automation through CPLT: count model/API usage, human review, corrections and failed attempts, then compare the same task at the same quality standard. Released time is capacity; cash is saved only when spending actually falls.

A chatbot can produce a fluent answer with a page number underneath it. The reviewer still needs to check that the passage supports the answer, the document is current and the person asking may access it. Include that checking time in the comparison.

What does a usable AI result cost?

A cheap model call can produce an expensive piece of work.

Someone still has to read the answer, open the references, catch the missing exception and fix the confident nonsense. If checking the answer takes longer than doing the work, I want that on the spreadsheet.

For a document workflow, I would track:

  • Time to an accepted result: from opening the task to finishing the checks and corrections.
  • Acceptance rate: how many results meet the agreed standard, including awkward cases.
  • Full running cost: model calls, hosting, storage, maintenance and human review.
  • Setup cost: integration, access controls, document preparation and training people to use it.

Count failed attempts too. Dividing costs only by the easy answers makes the spreadsheet look rather better than the business.

A practical measure is cost per accepted result: all the costs attributable to the test, divided by the number of outputs that pass the checks. Keep one-off setup costs visible separately when estimating ongoing costs.

State the test period, task volume and included cost categories before calculating the figure. Count each completed task once, even if it needed several attempts, and include the cost of rejected attempts in the numerator. If no results pass, there is no finite cost per accepted result to report.

An accepted result meets the written requirements after the agreed review; a fluent answer or a successful API request is not enough. Set those requirements before running the test, including when the system should decline or hand the task to a person. For a knowledge assistant, use the acceptance-check guide to specify the expected behaviour.

A fictional capacity calculation, with the assumptions showing

Here is a fictional example. These are round numbers to explain the calculation, not CPLT results or a savings forecast.

Suppose a team handles 200 document tasks a month. Each takes 20 minutes today. With AI, the complete task takes 12 minutes, including checking and corrections, at the same required quality.

That would release about 26.7 hours a month. At an assumed labour cost of €40 an hour, that capacity has an accounting value of roughly €1,067. Against assumed additional running costs of €400 a month, the difference is approximately €667 before setup costs; it is not cash saved or a payback forecast.

Now change the checked completion time to 18 minutes. The time value falls to roughly €267. The same €400 running cost wipes it out.

The difference in checked completion time changes the comparison. Neither scenario is a measured CPLT result.

Released time is capacity. You still need a credible use for it, such as clearing a backlog or taking on work. If staffing and other existing expenditure stay the same and the additional €400 is new expenditure, spending rises by €400; it does not fall by €667. Any cash saving needs an identified expense that actually falls. I would record the capacity plan and cash budget separately.

Start an AI pilot with one task you can judge

“Help the team with documents” is too vague to price or test properly.

“Extract these five fields, show where each came from, and flag anything missing” gives you something to check.

I would choose a recurring task with a named owner and a clear finish line. Then put together an authorised test set that includes incomplete files, conflicting versions and questions the documents cannot answer. Record how long the existing process takes before changing it.

Agree what failure looks like. A missing answer may be acceptable if it is clearly flagged. An invented answer presented as fact may make the whole workflow unsuitable.

Set the budget and the stopping point before the pilot begins. Otherwise there is always another model to try, another prompt to tweak, and another month to explain.

I explain how I start with one task in this video, using my AI likeness and voice; loading the YouTube player connects to Google.

Decide where your AI data may go

Before choosing a model, I would want a written answer to four questions:

  1. Which documents may enter the system?
  2. Which people may retrieve them?
  3. Which processing steps, if any, may use an external provider?
  4. What is stored in conversation history, logs and backups, and for how long?

Private AI can give you control over the infrastructure and the routes your data takes. That control still needs to be configured and tested. Putting a model on your own server does not, by itself, settle access permissions, retention or every other component's network behaviour.

If an external API is permitted, check the terms for that service and account. A consumer subscription is not an API budget, and a no-training commitment does not answer every question about retention.

The model choice comes inside those boundaries. A harder question is not permission to send more data elsewhere.

What should you own when the pilot ends?

I would expect a short, readable record: what was tested, what passed, what failed, what it cost and what remains uncertain.

If the decision is to continue, someone needs to own updates, access changes, backups and recovery. Your team should know how to export its data, replace a model and keep working when the service is unavailable. Those belong in the scope while everyone is still paying attention.

I describe how I scope AI assistance, automation and handover on the services page. There is also an interactive demo with fictional documents if you want something concrete to explore first.

For a concrete model-cost comparison, I checked Kimi K3 and GPT-5.6 Sol: cheaper tokens, a higher cost per task, using official pricing and independent benchmark data.

Bring me the task that keeps coming back

I build custom AI assistants and business automations through CPLT, with private hosting as an optional delivery choice. The useful starting point is one recurring task, the time it currently takes, and the data restrictions around it.

Book a free 45-minute remote scoping call. You receive a one-page written note on fit and next steps. Describe the task; keep confidential documents out of the initial message. Any paid work has a written scope and a price agreed first.

If the numbers do not justify building it, that is a useful result too.