Updated 13 September 2026: this article now presents a general method. Examples are illustrative, not customer results.
I separate the findings from the document that presents them. A model can prepare structured findings; a script can validate and format those findings; a person decides whether they are supported.
This is a design method, not a customer case study or a measured saving. Its usefulness depends on the task and on testing both the data and the rendered output.
The redesign: persist at assessment time, render by script
1. Keep a source record for each finding. Store the source reference, proposed conclusion and review status in a defined format. An assertion that appears only in a generated document is harder to trace back to its evidence.
2. Agree the allowed values. Where the workflow needs fixed categories or statuses, validate them against the agreed schema. Reject an unknown value rather than silently treating it as a new category.
3. Render from those records. A script can turn accepted records into a document without asking a model to rewrite the reasoning. This avoids an additional generation call; processing, storage and maintenance still have a cost.
4. Keep human decisions separate. A draft must not silently overwrite a person's approval or comment. Define which fields each component may change and test those permissions.
The rule that keeps it honest
If schema validation fails, inspect the source record. Patching only the exported document leaves two conflicting versions of the result, and a later export may discard the patch.
I also check the rendered document itself. Correct records do not prove that the renderer placed the right content, references and labels on the page. Data validation and visual inspection answer different questions.
What each layer is actually good at
- The model proposes a finding from the information it can access. That proposal still needs evidence and review.
- Code validates fields, applies ordering and formats the result using explicit rules.
- The person checks the source, resolves ambiguity and approves important decisions.
Stable data makes meaningful comparisons easier. Byte-identical files require additional controls: timestamps, metadata, ordering, fonts and renderer versions can change the bytes even when the visible content is the same. I test the level of reproducibility the task actually needs.
For a complete cost comparison, count generation, rendering, review, correction and maintenance. Fewer model calls alone do not establish a saving. See how to evaluate a paid AI pilot.
Questions people ask about deterministic deliverables
What goes wrong when a model writes the deliverable directly?
Regeneration can change wording and introduce unsupported additions. Preserve the underlying findings and check the document against them.
Why does reproducibility matter?
It helps reviewers identify what changed. Define whether the requirement is equivalent content, equivalent layout or identical bytes, then test that requirement.
What does the split look like?
Structured findings pass through validation and a renderer. A person retains control of review decisions.
What happens when the renderer rejects a row?
Correct or reject the source record, then render again and inspect the output.