A second model can check a draft, but a second answer is not automatically an independent opinion. If the reviewer sees the draft first, it may follow the same reasoning and miss the same problem.
I want the reviewer to read the source evidence first and form its own view. Only then should it compare that view with the draft.
The commercial question is straightforward: can a cheaper model do that job well enough to justify using it? I would test that question for the reviewing role, on an agreed set of examples. A good result would not establish that the same model can take over the drafting role.
Start with the error that matters
A reviewer can flag a sound finding unnecessarily, or approve a finding that is wrong. Both matter, but their consequences differ.
I would count those errors separately. A high agreement rate can hide false approvals, especially when most examples are easy. The test needs known mistakes, missing evidence and cases where the right answer is to stop and ask a person.
Keep the test independent
Before running the models, I would agree the expected answers and scoring rules with someone qualified to judge the work. The reference answer cannot simply be whatever the more expensive model said.
Both candidates should receive the information available to the reviewing role. Record the model version, settings, input and output so a surprising result can be investigated.
For a public demonstration, I would use synthetic examples or material that is explicitly cleared for publication. Confidential work does not belong in a published benchmark.
Measure the cost of review as well as the API bill
A cheaper API call may produce more work for the human reviewer. I would record false alarms, missed problems, unresolved cases, review time, latency and API cost.
The useful comparison is the cost of getting to an acceptable decision. Token price alone does not measure that.
A test result has limits
There is no universal number of cases that proves a reviewer is safe. The examples need to cover the intended work and its failure modes. Passing a small sample leaves considerable uncertainty; a failure can still reveal a concrete problem worth fixing.
I would keep a separate set of examples for the final check, repeat tests where outputs vary, and state what the evaluation did not cover. Any model change needs another check against the same acceptance criteria.
Keep the roles separate
If a candidate performs well as a reviewer, that supports using it as a reviewer within the tested scope. It does not establish that it can draft the work equally well.
When models disagree, I want that disagreement preserved for a person to resolve. The point of a second opinion is to make a missed problem more visible.