CAST 2026 - Open season: what a room full of sceptics taught me about the quality gate

I promised you this recap two weeks running. Here it is — and I'm glad I waited, because it took me a fortnight to work out what the conference was actually about for me, and it wasn't the talk I gave.

CAST 2026 - Open season: what a room full of sceptics taught me about the quality gate

I promised you this recap two weeks running. Here it is — and I'm glad I waited, because it took me a fortnight to work out what the conference was actually about for me, and it wasn't the talk I gave.

It was the twenty minutes after.

CAST doesn't let you finish, thank the room, and sit down. The format is adversarial by design: you present, and then the room takes you apart. Questions, counter-examples, "that's not what I've seen," people who've run the experiment you're describing and got a different result. If you've been reading this series, you already know why that stopped me in my tracks.

That's a review gate. A human one.

Open season is the quality gate, wearing a lanyard

Since the RPIQ post I've been circling the same idea: the Quality stage isn't a step you bolt on at the end, it's a gate — something a claim has to pass through, run by a reviewer whose whole job is to refuse to take the claim on trust. I've written about designing that gate into a model stack: separate reviewer, different vantage point, structured to catch what the builder is too close to see.

CAST's open season is that gate implemented in people. I stood up and made a confident set of claims about orchestration and model mixing. The room's job — the format's job — was to not accept those claims just because I sounded sure. And it didn't.

So this recap isn't "here's what I saw at a conference." Anyone can write that one, and it's the single post in this series anyone else could have written. This is the opposite: the room did to me exactly what a good reviewer does to a plan, and it advanced the argument this series has been making rather than pausing it.

Open season, mapped onto the RPIQ Quality stage. The room is the reviewer tier — its job is to refuse a confident claim without evidence.

What landed, and what flatly didn't

What landed: the orchestration itself. The core of the talk — mixing models, giving each tier the job it's actually good at, treating "which model" as a routing decision rather than a loyalty — got real engagement. People wanted the mechanics. The appetite in the room for multi-model setups was higher than I expected, which I'll come back to.

What didn't land: I am, openly, an AI enthusiast. I pitch the upside hard. And the sceptics in that room stayed sceptical. I didn't move them. A lot of work needed there, and I want to be honest about it rather than round it up to "great discussion."

Here's the part that took me two weeks to sit with, though: the sceptics weren't being unreasonable. That's the easy, wrong reading — that they just hadn't seen the light yet. The honest reading is that they'd been marketed at aggressively, promised things that didn't arrive, and had watched a lot of confident public claims about AI turn out to be noise — much of it from people who weren't building anything, amplified by clickbait and non-technical commentary.

Walk into that room as an enthusiast and you're carrying the field's credibility debt whether or not you personally ran it up. Being right isn't sufficient when the audience has been burned by confident claims before. The burden of proof is on me, not on them, and it should be.

Which closes the loop with the whole spine of the talk: the sceptics in that room were the quality gate. They did precisely what a good reviewer does — refused to accept a confident claim without evidence. That's the same lesson as the kill-gate in Post 8: confidence is not evidence, and a room that won't take your word for it is doing you a favour, even when it doesn't feel like one at the microphone.

It also justifies how I've been writing this whole series. The receipts. The real numbers. The honest retrospective on where the stack was overkill. The published admissions of getting things wrong. That's not modesty as a style choice — it's the only registration that works with an audience the field has already over-promised. You don't out-enthuse a burned room. You show your work.

Starlink start at Cocoa Beach

The two questions I couldn't answer well

The strongest material from the whole conference came from questions I didn't have good answers to. (This has a track record: the hallway question at CAST — "what do you actually lose without Fable?" — became Post 8 and carried it.) Two stuck.

One: how do you forecast any of this at all? Someone asked, reasonably, how I predict AI's growth and expansion when the ground moves this fast. And I don't have a good answer. Forecasting feels close to futile — the base rate of confident predictions in this space aging badly is very high, and I'm not going to pretend I've got a model that beats it. I think the honest position is that you build for optionality rather than forecast, but I said that with a lot less conviction than I'm typing it now.

Two — and this is the one that's still bothering me: is the token even the right unit of account? Someone pushed on token-based pricing, and the more I've thought about it the less convinced I am that it holds. It has real flaws. It isn't standardised — every vendor applies "tokens" inside its own model context, with only the rough base rules in common, so a token here and a token there aren't the same thing you can price against. It feels like transitional scaffolding, not a settled primitive.

That should worry me more than it worries most people, because the entire economics spine of this series is denominated in tokens. Post 3's cost-per-leverage argument. Post 9's "a model you have to ration can't hold the high-frequency seat." The whole case for why Fable didn't get its architect seat back rests on cost-per-call and rate limits. If the unit of account is wrong, that argument needs restating on firmer ground. It's rare to get handed, in public, a good reason to question your own foundation — so I'm going to take it. That question is getting its own post next week.

The room wanted more of this than I expected

The thing I underestimated going in was appetite. There's a genuine, growing hunger for orchestration and multi-model thinking — not from a fringe, but from a solid part of the room. People are past "does AI do anything useful" and into "how do I actually compose this into something reliable." That's a different conversation, and it's the one this series has been trying to have.

The room that ran open season. Most of the value of the trip was in the twenty minutes after I stopped talking.

Which points at something bigger than the talk. This is, I think, a once-in-a-career opening — for companies and for individual experts. The tooling is good enough to build real things and still raw enough that the patterns aren't settled, which means the people who put in the work now get to shape them. That window doesn't stay open. The cost of sitting it out isn't that you miss a trend; it's that you inherit patterns other people designed, on their terms, later.

Building on other people's work

Two talks changed what I'm doing, and I want to credit them by what I actually took, not by "great talk, very inspiring."

Matt Simons (Director of Engineering) gave a forward look at frontends and, more importantly for me, at agents becoming a commodity for ordinary people — not just technical experts. The democratisation angle. I've been building orchestration for people who already think in stacks and routing; his framing was a nudge that the interesting frontier is the non-expert. That's a gap in my own practice, and I'm taking it into the innovation work where I work — designing for people who will never read a post like this one.

Justin Bench (profile) runs a lot of experiments on how models actually behave, including self-hosted models. I use cloud models exclusively, so this is genuinely unexplored territory for me — and a pointed reminder that you can shape models, not only consume them. Training, fine-tuning, running your own — I've treated that as someone else's department. His work is why I'm now treating it as a gap in mine.

Both credits are useful precisely because each one names something missing in how I work — non-expert accessibility, and self-hosting / model shaping. That's worth more than applause. It's what changed.

Slides, and where this is going

The deck is here, pointed straight at where the series is heading — "Quality at Machine Speed," the book chapter this whole run has been building toward. The full CAST26 deck →

The through-line, if you've been following: the sceptics in that room, the reviewer tier in a model stack, and the kill-gate that decides whether a plan is worth building are all the same mechanism. Something whose job is to refuse a confident claim until it's earned. Quality at machine speed isn't about going faster through that gate. It's about building the gate so it still holds when everything upstream of it got fast.

If you're testing Fable 5 / Opus 5 too, where did it land for you? And the version I actually want answered this week: when has a room — or a reviewer, or a gate — refused to take your word for it, and been right to?


*Series: Post 1 — Fable 5 talks to machines better than to people · Show me the receipts · Show, don't tell · Why I took the architect tier out · The RPIQ loop · Where the stack was overkill · BowSmith case study · I lost my favourite tool · Opus Max as architect · Fable 5 came back