The arc ended on a stage: what a room of testers did to the thesis

The slides and the full talk breakdown live on the other track. This isn't that post. This is the other thing that happened at CAST26, and the reason this series ended on a stage at all: seven posts' worth of argument, finally standing in front of a room that got to push back.

The arc ended on a stage: what a room of testers did to the thesis

The slides and the full talk breakdown live on the other track — I put them there: CAST 2026: open season as a quality gate. This isn't that post. This is the other thing that happened at CAST26, and the reason this series ended on a stage at all: seven posts' worth of argument, finally standing in front of a room that got to push back.

For eighteen months this series asked one question out loud — could someone who had never written a line of Swift ship a production iOS app, with AI as the only developer? The whole arc is in one place here. CAST26 was that question being tested in public: not by me narrating the win, but by a room of testers and engineers deciding whether they actually bought it.

What the room did to the thesis

Three objections landed, in ascending order of how much they cost me.

The economics. Nobody disputed the ~€240/month. They disputed what it measured. The money was never the constraint — evenings were. The honest cost of BowSmith is attention, not subscription, and "attention" doesn't fit on a slide as cleanly as a euro figure does. Fair hit. The number is real; it's just answering a question almost nobody actually has.

The test count. I spent an entire post on the ~70,000-test suite. In the room, the useful version of that argument turned out to be the opposite of the post: test count is not a quality signal. AI writes tests that match the code, not the spec — so volume can be evidence of nothing. Evidence quality over evidence volume. If I were writing that post today I'd lead with the reframe and treat the number as a footnote, not a headline.

The one I still can't answer cleanly. The sharpest pushback wasn't about the app at all. It was: you keep saying "zero coding experience" — but you walked in with twenty years of QA judgment. Didn't that do the heavy lifting?

And, honestly — partly, yes. The talk's own conclusion is that the irreplaceable skill here isn't writing code; it's being the ground truth. Knowing the moment an agent is producing something plausible instead of something correct, and stopping it. Which means the series' framing — "a non-coder shipped this" — is true about the code and quietly misleading about the judgment. The code was the cheap part. I don't have a clean answer to "would this have worked without the QA background," and I'm not going to pretend I do. That's the loose thread the whole arc leaves hanging.

What I'd write differently

Eight weeks is long enough to see which posts aged worst.

The one about building a 17-agent team aged the hardest. That org chart was a beautiful, expensive detour; the move that actually worked was task-based, ephemeral agents, and I got there far later than I should have. If you read the arc, read that part as the mistake, not the method.

The method post I'd flip to an imperative: adopt RPIQ and write your CLAUDE.md on day one. Half of the early pain was the model rederiving the world every session because there was no constitution for it to stand on. The forcing function arrived eleven months in — everything before it was vibe-coding with extra steps.

And the quality post I'd cut the number from entirely and lead with the sentence that turned out to be the whole thesis: the quality role didn't disappear, it moved — from reviewer to architect of evidence. That's the line I'll be expanding in the "Quality at Machine Speed" chapter, whenever that book lands.

Closing the loop

That's the arc closed. Seven posts asked whether AI could stand in for a feature team. CAST26 was a room of testers deciding whether the answer held. BowFest was the same question asked by archers instead of engineers. Two audiences, one claim, tested from both ends — and the most useful feedback came from the people best equipped to tell me I was wrong.

If you want the whole thing in order, it's all here. Missed the talk? The deck and breakdown are here.


Series — From Vision to Production: the full arc, in order · CAST26, the deck & the talk · BowFest 2026, the product in the field