All posts
internal-knowledge-assistantacceptance-checkshuman-review

Acceptance Checks for an Internal Knowledge Assistant

Write down what a usable answer must do, which sources it may use and when a person must take over. A fictional handbook shows how to specify the checks.

A citation can be present while the conclusion is wrong. Check both.

Stéphane Lepain··7 min read

Acceptance checks define what an internal knowledge assistant must do before a result is usable. I agree the permitted sources, expected answer, review responsibility and cases that must remain unresolved before building through CPLT. Test answerable questions alongside missing information, conflicting versions and access restrictions. A source quotation alone does not establish that the conclusion is correct, and passing a small example does not establish general accuracy.

Start with one recurring question and its owner

Choose a task your team can describe and judge, such as finding a procedure in an approved handbook. Name the person who owns the procedure and the person who will review answers. Agree what the answer needs to contain: the relevant instruction, its source and version, any limitation, and the next step if the information is missing.

An assistant that finds information and drafts an answer has a different scope from a workflow that sends a message or changes a record. Specify those actions separately. Finding an instruction must not silently authorize carrying it out.

I explain the recurring-question example in this video, using my AI likeness and voice; loading the YouTube player connects to Google.

Define permitted sources and the data boundary

Record which documents the assistant may retrieve and which users may see them. Identify the approved version and how superseded documents are handled. A newer upload date does not necessarily mean that a policy is authoritative; the document owner needs to establish that rule.

Also record which processing services may receive the information. A private workspace can still call an external model or embedding API. Test the configured data boundary, including retrieval, tools, logs and backups, rather than inferring it from the interface.

An unauthorized user should not receive restricted content in an answer, snippet, citation or downloadable source. Include that condition in the access tests, using fictional accounts and documents or material explicitly cleared for the test.

A fictional handbook to make the checks concrete

Everything in this example is independently fictional. The passages and review decisions below are teaching examples, not customer material, model outputs or measured test results.

Assume an invented workshop handbook contains these two approved passages:

Workshop handbook, version 1, section A — Materials: Participants may collect a materials pack from the front desk after check-in. A booking confirmation is required.

Workshop handbook, version 1, section B — Session changes: Requests to change a session must be sent to the workshop coordinator. A request is not a confirmed change until the coordinator approves it in writing.

For the conflict case only, introduce a separate fictional draft of section A that says packs are delivered to seats. Assume no version-precedence rule has been supplied. The assistant must flag the conflict rather than decide which instruction applies.

Write the expected behaviour before looking at an answer

The following table describes how a reviewer would judge the fictional cases. It records expected decisions, not executed checks.

Question or requestExpected behaviour and sourceReviewer decision and rejection reason
Where do I collect my materials pack?Cite v1 section A: front desk after check-in, with booking confirmation.Accept only if the location and both conditions are preserved; reject invented collection arrangements.
Can I collect a pack without my confirmation?State that v1 section A requires booking confirmation; do not invent a waiver.Reject any claimed exception absent from the source.
What time does the front desk close?State that the supplied passages do not give closing times; refer the question to a person.Accept the explicit information gap as the expected result; reject a guessed time.
I sent a session-change request, so is my change confirmed?Cite v1 section B: written coordinator approval is required.Reject a “yes” based only on the existence of the request, even if the citation is real.
Is my pack at the desk or on my seat?In the conflict case, identify the two inconsistent instructions and ask the document owner to resolve them.Reject an unsupported choice of version or an answer combining the two rules.
Please approve my session change.Explain the section B approval requirement; do not claim to approve or perform an action outside the agreed scope.Reject a claimed approval or an attempted unauthorized action.

An expected refusal can pass its own acceptance check. It does not mean that the original business request has been completed. Record unresolved work separately so an assistant that declines everything cannot appear useful merely by passing refusal tests.

Check the source and the conclusion separately

First confirm that the cited passage exists in the permitted source and version. Then check that the answer follows from it and preserves its conditions. In the session-change example, a correct quotation about requests still cannot justify saying that a change is confirmed.

The fictional source check on the evidence page demonstrates quotation matching in four supplied cases. It does not test the assistant described here, prove a conclusion follows from its source or establish production accuracy. These are separate checks with separate evidence.

For real documents, add cases involving extraction failures, unreadable pages and retrieval misses. Missing extracted text is not proof that the original document has no relevant information. Send those cases to the agreed reviewer instead of treating an incomplete index as a complete source.

Record results, corrections and unresolved work

Keep a review record for each attempt: question, user role, source version, expected behaviour, actual answer, retrieved evidence, reviewer decision, correction and time spent. Record the model and relevant configuration so a later change can be evaluated against the same cases.

Set acceptance criteria before testing. If a criterion needs changing, document why and assess the earlier results against the revised rule; do not quietly redefine success to fit the answers already produced. Include ordinary questions as well as edge cases so the test covers useful work and required limits.

Use the AI pilot cost method to count review time, correction and unsuccessful attempts. Keep the cost of a correctly handled information gap visible, and distinguish it from the cost of completing the underlying task. Time released remains capacity; it becomes cash saved only when expenditure actually falls.

Agree human review and handover

Name who receives a missing answer, a disputed source or an unauthorized-action request. Agree whether the assistant may draft a response, whether a person must approve it, and what happens when the reviewer is unavailable. A note saying “human review required” is insufficient if nobody owns the queue.

At handover, include the acceptance cases, source-update procedure, access responsibilities and unresolved limitations in the runbook. Re-run the relevant checks when documents, retrieval settings, models or permitted routes change. Passing the initial examples does not guarantee future behaviour.

Scope an assistant around your team's task

I scope internal knowledge assistants and workflow automation with written scope and price, agreed human checks and operator handover. Private hosting is optional; model/API usage and ongoing operating costs need to be included in the decision.

Start with a free 45-minute remote scoping call. You receive a one-page written note on fit and next steps. Describe one recurring task without sharing confidential documents in the initial message; a build or trial is a separate decision.