Opus 5: the middle ground I actually chose

A tool I relied on vanished, and I scrambled. I wrote a job description for the empty seat and dropped a stand-in into it. The favourite came back, and I found it a different seat rather than its old one. Then I went to a conference and let a room full of sceptics pressure-test the whole thing.

Opus 5: the middle ground I actually chose

A tool I relied on vanished, and I scrambled (Post 7). I wrote a job description for the empty seat and dropped a stand-in into it (Post 8). The favourite came back, and I found it a different seat rather than its old one (Post 9). Then I went to a conference and let a room full of sceptics pressure-test the whole thing (Post 10).

Every one of those was forced. Something changed underneath me and I responded. That's fine — most real systems are shaped more by what breaks than by what anyone planned. But it does mean I've never actually written down the moment where the stack got chosen rather than patched.

This is that post. The architect seat is settled now, and it's held by Opus 5 — not because anything broke, but because, with all three models available and nothing on fire, it's the right occupant.

The three candidates, on the table at once

Here's the thing that never happened during the outage: a fair comparison. When Fable 5 was gone, "which model should hold architecture?" had a trivial answer — whichever one was still working. When it came back, I'd already reorganised around its absence, so the return was a reshuffle, not an audition.

For the first time, all three were genuinely on the table for the same seat: Fable 5, Opus 5, and Opus. And the seat has a real job description now, because I was forced to write one in Post 8: turn fuzzy intent into a structured plan, hold boundaries, decide whether a plan is even worth building, decompose it into work the workers can execute. That seat gets invoked dozens of times a day, and every invocation blocks something downstream.

So I scored the three the way you'd score any hire — not on raw talent, but on fit for this role, with its particular shape of demand.

The architect seat scored on three axes at once — capability, cost per call, and availability at high frequency. Opus 5 isn't top on any single axis. It's the only one that clears the bar on all three.

Capability. Fable 5 has the highest ceiling — it reframes problems, spots the second-order consequence, sees the shape you didn't ask for. Opus is close behind and rock-solid. Opus 5 sits between them: not the sharpest planner I've used, but comfortably past the bar the seat actually needs. Architecture rewards reliable good judgement invoked constantly more than it rewards brilliance invoked occasionally.

Cost per call. This is where the ceiling stops mattering. The architect seat is the highest-frequency seat in the stack, so its cost is multiplied by every task, every re-plan, every gate. Fable 5 is expensive per call; Opus is middling; Opus 5 is cheap enough that I never think twice about invoking it. On a seat you hit a hundred times a day, "cheap enough not to think about" is a capability of its own.

Availability at frequency. The quieter axis, and the one that actually decided it. A model behind tight rate limits can't hold a seat you lean on constantly — the moment you're near a limit, you start batching decisions that shouldn't be batched, or quietly skipping the gate. Opus 5 has the headroom to be invoked freely. Fable 5 doesn't, which is exactly why it moved upstream to ideation instead of back into the loop.

Score it across all three and Opus 5 wins — not by being the best model, but by being the best fit for a high-frequency, blocking, judgement-heavy seat. The middle ground isn't a compromise here. It's the target.

What changed in the hand-offs

Choosing Opus 5 deliberately did something the outage never could: it let me collapse two seats into one on purpose.

During the stopgap, architecture and coordination were still two conceptual jobs I happened to route to the same model. Now they're genuinely one seat. Opus 5 turns intent into a plan and holds the plan together as the workers execute against it — the boundary-keeping, the "does this sub-task still serve the goal," the first quality gate. That used to be a hand-off. Now it's one continuous context.

Two things got measurably simpler downstream.

First, context stopped leaking at the plan→coordinate boundary, because there's no longer a boundary there to leak across. The most expensive failure mode in a multi-tier stack is the architect deciding one thing and the coordinator subtly re-deciding it. Collapsing the seat removes the gap where that happens.

Second, the workers got cleaner briefs. Sonnet is unchanged — it still writes the code — but it's now taking instructions from a single coherent voice rather than from a plan that one model wrote and another model re-interpreted. Fewer "wait, which instruction wins" moments. Less rework.

The difference between a stack you patch under pressure and one you choose on merit. Same three models, same three seats — but the second one was designed, and it shows in where the context flows.

None of that is a Opus-5-specific miracle. It's what happens when you get to make the seat assignment as a decision instead of inheriting it from a crisis.

Why this is the interesting post, not the boring one

An enablement post — "here's the model I settled on and why" — reads like the least dramatic entry in a series full of things going wrong. I'd argue it's the point the whole series was building toward.

The thesis from Post 8 was design the seat, not a favourite. Posts 7 through 10 tested that thesis under duress — could I define a seat, fill it with a stand-in, and reorganise around a returning favourite without falling apart? Yes. But duress is a low bar. Anything holds together when the alternative is not working at all.

The real test of "design the seat, not a favourite" is what you do when nothing is forcing your hand. When your favourite is available, your old setup is working fine, and you could just leave it — do you still make the assignment on merit? This post is me answering yes. I looked at three available models against a written job description and chose the one that fit the role, not the one I'm fondest of (that's Fable 5, and it's upstream) and not the one with the highest ceiling (also Fable 5).

That's the difference between a stack that got patched and one that got designed. A patched stack is a history of what was working the day each thing broke. A designed stack is a set of deliberate matches between seats and occupants, each defensible on its own terms. Mine is finally the second kind — and it took the outage, the return, and the conference to get me there.

Where the quality gate lives inside this — the first-pass review the architect seat runs before anything reaches the workers — is the thread I'm pulling into the forthcoming book chapter on "Quality at Machine Speed." This post is the seat assignment; that chapter is what the seat is for.

Next week

The economics underneath all of this — cost per call, cost per unit of leverage, the whole argument for why an affordable model beats a brilliant one in the high-frequency seat — is denominated in tokens. A question from the CAST26 floor that I couldn't answer well has been bothering me since: is the token even the right unit of account? That's next week, and it's a sharp one, because it questions the foundation the last several posts have been standing on.

If you're testing Fable 5 / Opus 5 too, where did it land for you? And the deeper version: when did you last assign a seat in your stack on merit, with everything available and nothing forcing your hand — versus inheriting the setup you happened to have the day something broke?


Series: Post 1 — Fable 5 talks to machines better than to people · Show me the receipts · Show, don't tell · Why I took the architect tier out · The RPIQ loop · Where the stack was overkill · BowSmith case study · I lost my favourite tool · Opus Max as architect · Fable 5 came back · CAST26: open season as a quality gate ·