Reviewer design: catching what the architect tier misses

A model's output is shaped by its assumptions, it brings the identical set of assumptions to the check. So it reads its own work and finds it sound, not because the work is sound, but because author and reviewer share a blind spot in exactly the same place.

Reviewer design: catching what the architect tier misses

Last week I argued that delegation isn't done until you've decided who checks the work — that "who checks whom" belongs in the brief, not bolted on afterwards. A few people agreed in principle and then asked the obvious follow-up: fine, but who should check it, and how? Because the default answer most stacks reach for is the worst one available.

The default is: let the thing that did the work check the work. And that isn't a review. It's a signature.

Same-tier review is a rubber stamp

Here's the mechanism, because it's worth being precise about why self-review fails rather than just asserting it does.

A model's output is shaped by its assumptions — what it treated as obvious, what it didn't think to question, where it rounded off. When the same model reviews that output, it brings the identical set of assumptions to the check. So it reads its own work and finds it sound, not because the work is sound, but because author and reviewer share a blind spot in exactly the same place. The reviewer is most confident precisely where both of them are wrong together.

This is the uncomfortable part: a confident wrong answer and a confident review of that answer are produced by the same machinery. Asking a model "are you sure?" and getting "yes" back is not evidence. It's the same draw from the same distribution, dressed as a second opinion. You haven't added a check. You've added a co-signer.

Human teams learned this the hard way and built around it — that's what separation of duties, second readers, and external audit all are. The principle ports straight across: a reviewer that shares the author's blind spots isn't a check. The whole value of a reviewer is that it fails differently from the author. Independence is not a nice-to-have property of the review. Independence is the entire product.

What a reviewer is actually for

Say it plainly, because it reframes the whole job: the reviewer's purpose is not to re-do the work and see if it matches. It's to disagree with the work from a different vantage point — and to surface the thing the author structurally could not see.

That reframes what a good review even looks like. "Does this look fine?" is the wrong question, because a same-context reviewer will almost always answer yes. The useful question is adversarial: where does this break? what did the author assume that isn't true? what's the case they didn't consider? A reviewer briefed to find the failure will find more than one briefed to bless the output — not because it's smarter, but because it's looking for a different thing.

Which means review has design parameters, and they're all about distance from the author:

  • Different tier — a reviewer at a different capability level reasons about the work differently than the author did.
  • Different context — a reviewer that never saw the author's chain of reasoning can't inherit its assumptions; it has to re-derive, and that's where gaps show.
  • Different framing — brief the reviewer to break the thing, not to approve it.

The more of those you stack, the more independent the check, and the more it actually catches. A review that shares the author's tier, context, and framing is independent on zero axes. It's the rubber stamp again, wearing a lanyard.

Same-tier review (author's blind spot passes straight through — a rubber stamp) vs. cross-tier review (an independent reviewer, different tier/context, adversarial brief, catches the miss). Below: who checks whom, mapped across architect / orchestrator / workers.

The architect tier isn't exempt

Now the part that's easy to skip, because it's the one that implicates the top of the stack.

It's tempting to think reviewer design is a problem for the workers — the execution tier produces code, something checks the code, done. But the most expensive errors in an orchestrated stack don't start at the bottom. They start at the top, in the plan. The architect tier sets the approach, the decomposition, the constraints everything downstream inherits. And it plans with its own assumptions baked in — the same way any author does.

A plan reviewed by the planner inherits every one of those assumptions unexamined. If the architect framed the problem slightly wrong, a worker will execute the slightly-wrong frame flawlessly, and the review at the bottom will confirm the code does exactly what the plan said — while the plan was the thing that was off. You passed every local check and shipped the wrong thing, confidently, with a clean audit trail.

So the uncomfortable rule is that the higher the tier, the more load-bearing its unchecked assumptions become — and therefore the more it needs an independent check, not less. The architect tier is where review is most often skipped (it's the smartest thing in the room; who reviews it?) and where skipping it costs the most. Reviewer design has to answer "who checks the architect?" first, not last.

Designing the review into the stack

Put together, this is a design problem with a few concrete moves, and none of them are exotic:

  • Route the review across tiers, not within one. The check for a piece of work comes from somewhere that doesn't share the author's failure modes. Cross-tier is the cheap version of independence.
  • Strip the context for the reviewer. Don't hand the reviewer the author's reasoning. Give it the output and the acceptance bar and let it judge cold; if it has to re-derive and lands somewhere else, you've found something.
  • Brief the reviewer adversarially. "Find what's wrong with this" surfaces more than "check this." Make breaking it the job.
  • Gate the plan, not just the output. Review the architect's plan before implementation starts — the cheapest place to catch a wrong frame is before anyone's built on it. A gate at the end only catches execution errors; a gate at the plan catches the expensive ones.

Notice that every one of these is a routing decision. Which is the thread from two weeks ago closing: routing isn't only about who does the work, it's about who checks it, and where the gate sits. Review is a first-class citizen of the stack's design, not a step you tack on when something's gone wrong.

The book-chapter version of this

This is the ground the forthcoming book chapter on "Quality at Machine Speed" is built on, and it's worth stating the thesis the chapter takes apart: quality at speed isn't an inspection you run at the end. It's a property of how the loop is wired — specifically, of how independent your checks are from the things they're checking. Bolt-on QA scales badly because it's always one reviewer against a rising tide of output. Designed-in review scales, because independence is structural: it's in the routing, not in how hard someone squints at the end.

I'll keep the reference light until the chapter's date is confirmed — but this post is the foundation it stands on. Reviewer independence is the load-bearing idea.

The short version

If the thing that wrote it also checks it, you don't have a review — you have a co-signer with a shared blind spot. The reviewer's whole job is to fail differently from the author, so design the check for distance: different tier, stripped context, adversarial framing. And don't exempt the top of the stack — the architect's unchecked assumptions are the most load-bearing and the most expensive, so gate the plan before you gate the output. Review isn't a step. It's a routing decision you make on purpose.

If you're testing Fable 5 / Opus 5 too, where did it land for you? And the reviewer version: who checks your top tier right now — and honestly, how would you know if it was confidently wrong?

Next week, the sharp flip side of all this care: the overengineering tax — when three tiers (and three reviews) cost more than one, and how to tell the difference.


Diagram: same-tier review (rubber stamp) vs. cross-tier review (independent check), with the who-checks-whom mapping across the stack.

Series: Post 1 — Fable 5 talks to machines better than to people · Show me the receipts · Show, don't tell · Why I took the architect tier out · The RPIQ loop · Where the stack was overkill · BowSmith case study · I lost my favourite tool · Opus Max as architect · Fable 5 came back · CAST26: open season as a quality gate · Opus 5: the middle ground I chose · Is the token even the right unit of account? · Routing rules: which tier gets which task · Stop prompting, start delegating ·