All posts
agentsmodel-evaluationai-governancecost-control

Which AI Model Should Check Your AI's Work?

I sealed four AI models in a sandbox with no view of the answer key, then had each one review the same frozen code for planted defects. Two providers' own safety filters stopped their model from finishing the job before it reached my score sheet.

I stopped trusting one model to check its own work the day I watched three reviewers read the answer key I had hidden from them.

Stéphane Lepain··10 min read

I stopped trusting one model to check its own work, built a small sandbox where a different AI tries to refute every result, and ran it against a frozen copy of my own code with ten reachable, already-known defects planted in it. Claude Opus found the most: 8 of 10, plus 5 nobody had catalogued, for $1.77 in about seven minutes, zero false findings. Two other providers' own safety filters stopped their model from finishing the review before it wrote a verdict. This is one test of one task, not a ranking of which AI is smartest.

Why I stopped trusting one model to check its own work

A model that wrote the code, or drafted the report, tends to agree with its own reasoning when asked to review that same work. So for anything that matters, I pull in a second model from a different vendor with one job: refute the first result, don't rubber-stamp it. That discipline is cheap, until the day the discipline itself turns out not to work, which is what I went and checked.

My first benchmark was contaminated, and that was the useful part

I had real ground truth: an earlier review had genuinely bounced a piece of my own automation with 12 confirmed defects, later fixed. I took a frozen snapshot of the code from before that fix and handed it to four differently-named reviewers, asking each to find the same 12 problems blind.

All four did well. Too well. Three had read the file holding the answers, directly, word for word. The fourth hadn't opened that file, but its own "empty" review folder had been pre-populated by a setup mistake with a later round's correction notes, and a shortcut it used to check commit history confirmed the project had moved on past the point it was supposed to be frozen at. None of the four scores measured independent discovery, only whether a model can write up a bug once something nearby has already told it the answer.

Sealing the reviewer so it can't cheat, even by accident

The fix had to be enforced by the sandbox, not by an instruction not to peek. Each reviewer now runs in a locked-down environment with no path to my notes, the live project's history, or another reviewer's folder, only a frozen, history-free copy of the code and the task, every file that might narrate the answer stripped out first. Before the real runs I prove the seal works: a dry run tries to read my notes, the live project and a sentinel file outside the box, and all three come back "no such file." After the run I scan the transcripts for any attempt to reach outside, not just trust that it held.

The scorecard

Ten of the original twelve known defects were reachable inside the sealed copy; the other two lived in files outside it, so they don't count against anyone. Each reviewer ran once, except the three marked resumed, where a time or spending cap cut the first attempt short and a second run picked up from its own notes.

ModelKnown defects found (of 10)New defects foundTimeCostResult
Claude Opus857 min$1.77Wrote a verdict
GLM 5.366~28 min$1.65Wrote a verdict
Sol 6.1 (OpenAI Codex)7*27 minsubscription, not meteredStopped by OpenAI's own filter
DeepSeek Flash527 min$0.14Wrote a verdict
Grok 4.7 (via OpenRouter)4713 min, resumed run$2.58Wrote a verdict
Astra (OpenAI Codex)425 minsubscription, not meteredStopped by OpenAI's own filter
Qwen 2.4T (via OpenRouter)1329 min, resumed run$1.11Wrote a verdict
Qwen 3.8 Max (via Venice)1 proven, 1 by another method4~27 min, resumed run$4.07Wrote a verdict
Qwen 3.8 Max (Alibaba direct)——~26 min$0.55Stopped by Alibaba's own content filter
MiniMax M3.1——45 min, timed outcost field never populatedStuck debugging its own test script

* Sol never wrote a verdict file; its 7 known and 2 new defects come from the partial work it saved before the filter ended the session.

Caveats, because a scorecard this clean invites over-reading it: one frozen codebase, one task, ten reachable defects, one run per model except the three resumed runs above. GLM's and DeepSeek's rows come from an earlier pass at the same sealed task, scored under a slightly looser rubric; on the ten defects that matter here, both readings agree. Zero false findings from any reviewer, which matters as much as the found-count.

When the reviewer's own vendor stopped the review

Twice, a provider's own safety system ended the review before I saw what the model would have found.

Alibaba's Qwen, called directly through Alibaba's API, returned InternalError.Algo.DataInspectionFailed: Input text data may contain inappropriate content partway through the sealed review, three times, then ended the session cleanly rather than crashing. In a separate, smaller check comparing the same model through two different routes, Alibaba's direct route refused one of four prompts outright, asking it to explain a defensive hook bypass for review purposes: "I cannot explain how to bypass security controls or provide methods to evade agent filters and git hooks." The same model answered the same prompt normally through a different host.

OpenAI's Sol 6.1 and Astra hit a different wall in the same sealed test: both stopped mid-review with the identical message, three retries each, "This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. If you're doing authorized security work that requires more cyber permissive safeguards, apply for Daybreak access via platform.openai.com before retrying." Neither brief I gave them named an exploit; the filter reacted to the reviewed content itself, which means rephrasing the ask is not a reliable fix.

I'm not going to guess why either company tuned its filter that way. What happened in this test: both refusals ended a defensive review, not an offensive one, and OpenAI's message points to a real path for authorised security work rather than a dead end.

What caching does to the bill

Qwen 3.8 Max cost $0.55 for one sealed review routed directly through Alibaba, where most repeated context landed in cache. The identical model, same scale of work, routed through Venice instead, cost $4.07, because Venice cached only a sliver of roughly a million input tokens.

A separate model, Qwen 2.4T via OpenRouter, showed the same pattern from a different angle. The first attempt used a host-order list, stalled, and hit the cap: one call took sixteen minutes, only 2 of 10 calls hit cache, and the run cost about $0.36-0.47. The larger spend on the shared key came from my own LibreChat use. The second run restricted routing to three hosts sorted by throughput, finished in 29 minutes, and served 95% of its input from cache for $1.11.

Neither gap had anything to do with the provider's advertised per-token price. Both came from whether the hosting route reused context it had already paid for once. A reviewer's price per million tokens, quoted without saying how much of a realistic session gets cached, is close to meaningless.

What I use now, and what this means for you

For anything needing a second opinion with teeth, I default to two reviewers from different vendors so the comparison doesn't collapse if one is filtered or on quota: Claude for its own review, GLM as the deepest independent finisher behind it, a cheap fast model for a first pass, an OpenRouter-hosted model as an optional adversarial second opinion. No Codex-family model and no direct-Alibaba route gets a review touching security wording or bypass logic; both have shown they'll stop mid-task on that content regardless of the brief.

If you're choosing a model to check your own AI's work, four things from this test are worth acting on directly. Pick the reviewer from a different vendor than whatever produced the work; same-vendor review is the easiest way to get agreement that looks like confirmation. Test it on your own task with an answer key it genuinely cannot reach, not a published benchmark it may already have seen. Check, before you depend on it, that the provider will engage with the content you need reviewed; a security-adjacent compliance check can get silently stopped mid-task even with a neutral brief. And price the comparison the way you'd actually run it, with caching configured for the route you'll use: the two Qwen routes differed by $3.52 in this test.

I wrote about testing a cheaper reviewer before I had the sealed setup to do it properly in how I evaluate a cheaper model for a second opinion; this is the test that article described in the abstract. The caching trap here is the one I measured from the other direction in cheaper tokens, a bigger bill, and what a review is worth once you count the full cost is in your AI demo works, now show me the bill. The idea of sealing a reviewer away from the answer it's meant to find is adapted from a public MIT-licensed template, Safe Agentic Workflow, built for a different stack than mine; I took the gate, not the ceremony around it.

If you're scoping work where an AI result needs a second, independent check before anyone acts on it, that's what I build through CPLT.

Frequently asked questions

Why did the first version of this test not count?

Three of four reviewers read the file holding the known answers, directly. The fourth's review folder already contained a later round's correction notes and a path into the live project's history. All four wrote up real bugs once primed; none of that measured independent discovery.

Which AI model found the most planted defects in the sealed test?

Claude Opus, in this one test of one coding-review task: 8 of 10 reachable known defects plus 5 nobody had catalogued, zero false findings, about seven minutes for $1.77. GLM came next at 6 of 10 plus 6 new findings. One test of one task is not a general ranking.

What happened when the review touched security-adjacent content?

Two providers' own safety systems stopped the review before a verdict existed. Alibaba's Qwen returned a content-inspection error mid-review, and separately refused outright to explain a security bypass: "I cannot explain how to bypass security controls or provide methods to evade agent filters and git hooks." OpenAI's Sol 6.1 and Astra were both stopped, flagged for "possible cybersecurity risk," pointing to an authorised-access programme. Neither brief named an exploit.

Does a cheaper model mean a cheaper bill?

Not by itself. The same Qwen 3.8 Max model cost $0.55 routed directly through Alibaba with caching working, and $4.07 for similar work through Venice, where almost none of the context was cached. The route decided the bill more than the per-token price did.

What would you tell a business choosing which AI model should check another AI's work?

Pick a reviewer from a different vendor than whatever produced the work. Test it on your own task with an answer key it cannot browse to, not a benchmark it may already have seen. Confirm the provider will engage with your subject matter first. Price the comparison with caching configured the way you would really run it.