The Complete AI Orchestration Journey: From Claude Web Chat to Production iOS App
The reference notes I wrote in November 2025, before BowSmith shipped: four phases, the original 17-agent team, the prompt template that cut context drift, and the failures. Lightly updated, with pointers to where the thinking went next.
Part 2: All the materials, lessons, and patterns from 18 months of AI-assisted development
A note before you start. I wrote this in November 2025, as the companion "vault" to my first talk about BowSmith, and then never hit publish. Since then the app shipped to the App Store (July 2026), the 17-agent team was replaced, and the stack grew a third tier. I've kept the original notes because they are the honest record of how it started, and I've marked where later posts take it further. If you want the polished version, start with From Vision to Production: the whole series, in one place. If you want the raw notebook, read on.
In Part 1, I shared the overview of my experiment – building BowSmith, an iOS archery companion app, using AI as my primary development resource. Now it is time to open the vault.
This post contains everything: the presentation materials, the detailed phase breakdowns, the agent structures, the prompt templates, and the honest failures. Use it as a reference, a learning resource, or a starting point for Your own AI-assisted development journey.

Table of Contents
- The Central Question
- Phase 1: Claude Web Chat
- Phase 2: Cursor IDE Integration
- Phase 3: Claude Code + Agents
- Phase 4: Claude Code + Skills
- The Agent Team Structure
- Prompt Templates That Work
- What Worked vs What Failed
- The Human Role Evolution
- Lessons for Testers
- Presentation Materials
- Resources and Links
The Central Question
"Can one person with domain expertise orchestrate AI systems to build a production-quality mobile app? What are the limits? What is the process?"
My constraints:
- Full-time job (no 40hr/week development available)
- Solo founder (no team)
- Limited budget (bootstrapped)
- 18-month timeline to alpha
- Zero iOS/Swift experience
- 20 years software testing background
The product:
- BowSmith – iOS archery companion app
- Target: Compound bow enthusiasts in 3D and field archery
- Core features: Competition tracking, bow tuning documentation, gear management, progress analytics
- Market: Initially US (underserved compared to Olympic/recurve apps)
Phase 1: Claude Web Chat (Feb-Apr 2024)
I started where most people start: a chat window. A basic Next.js landing page, feature ideation, requirements, the first React components. Everything moved by copy and paste.
Tools & Approach
- Claude Opus/Sonnet for conversations
- Manual copy-paste of code
- Iterative prompting and refinement
- No persistent context or workflow
What Worked ✅
- Rapid prototyping of ideas
- Clear explanations of archery domain concepts
- Quick UI/UX mockups
- Validation of the AI-assisted development concept
- AI understood domain-specific problems without being trained on them
What Did Not Work ❌
- Context loss between sessions
- Manual file management nightmare
- No code execution or testing
- Repetitive explanations of project structure
- Copy-paste workflow does not scale
Key Lesson
"Quality of output directly correlates with quality of prompting. But copy-paste workflows hit a ceiling fast."
Phase 2: Cursor IDE Integration (May-Sep 2024)
Moving into an AI-integrated IDE was the first real step change. The model could see the codebase, edit in place, and I could run a suggestion seconds after reading it. This is where actual Swift work began: session tracking, gear management, competition scoring, CoreMotion sensor data, the first SwiftUI architecture and an Appium automation setup.
What Worked ✅
- AI could see and understand full codebase
- Instant code modifications in place
- Reduced context switching
- Could run and test suggestions immediately
- Faster iteration cycles
- Model improvements noticeably better output
What Did Not Work ❌
- Single long conversations became confused
- Hard to manage multiple parallel tasks
- AI would "forget" architectural decisions
- No specialization – one AI doing everything
- Context window limitations with large codebases
Key Lesson
"Integration with development environment is game-changing. But single-agent approach does not scale to complex projects."
Phase 3: Claude Code + Agents (Oct-Dec 2024)
This is where I stopped treating AI as an assistant and started treating it as a team. I moved to Claude Code and built 17 specialised agents across five layers:
- Business: product strategist, market analyst
- Design: UX designer, SwiftUI developer
- Development: feature developer, data architect, API integrator
- Quality: code reviewer, test automation, accessibility auditor
- Operations: documentation writer, release manager
I also split the work into two repositories: the app, and the UI automation that verifies it.
What worked: clear separation of concerns, focused output, parallel streams, and noticeably better quality once review was a separate role.
What didn't: agents didn't share context across repositories, so I became the courier. The setup overhead was real. And several agents barely ever ran.
The lesson: specialisation matters, but start with fewer agents and add them when the work asks for them.
Where this went next: the role-based team didn't survive contact with 2026. I replaced it with task-based Agent Teams, and wrote up why in Building the team: from 17 role-based agents to task-based Agent Teams.
Phase 4: Claude Code + Skills (Jan-Nov 2025)
By late 2025 the system looked like this: skills for consistent artefacts, a project-wide rules file for persistent context, a universal prompt template, parallel streams for app and test automation, and different models for different jobs (the strongest one for reasoning, a faster one for implementation).
Context management still needed attention. Long sessions still needed resets. Some calls were still better made by a human. But development sessions could now stay coherent across complex features, which a year earlier they could not.
Where this went next: "different models for different jobs" became the three-tier stack (architect → orchestrator → workers) and the RPIQ loop. Start with RPI → RPIQ: the method that made AI build production software.
The prompt template that cut context drift
This is the single artefact from the period I'd still hand to anyone starting out. In my notes at the time it reduced context drift by roughly 80%. That was my own rough measurement, not a controlled study.
## Context
[Project name, current phase, relevant background]
## Current State
[What exists, what changed recently, relevant file paths]
## Task
[Specific, measurable objective]
## Constraints
[Technical limits, style rules, things to avoid]
## Expected Output
[Format, structure, deliverables]
A real example, from the scoring feature:
## Context
BowSmith iOS app, implementing competition scoring.
## Current State
- CompetitionView.swift exists with basic layout
- ScoreEntry model defined in Models/
- No scoring logic yet
## Task
Implement 3D archery scoring:
- Standard 3D rings (11, 10, 8, 5, 0)
- Running total
- 14-target rounds
## Constraints
- Follow existing SwiftUI patterns
- Use @Observable (iOS 17+)
- Unit tests for the scoring calculation
## Expected Output
- Updated ScoreEntry model
- ScoringService with calculation logic
- Tests in Tests/ScoringTests.swift
The template isn't clever. It works because it forces me to know the current state before I ask for the next one.

What worked, and what failed
Worked:
- Specialised roles with clear responsibilities
- The universal prompt template
- Two repositories: the product, and the thing that checks the product
- Different models for different kinds of thinking
- Domain expertise driving decisions: AI implements, the human validates
- Adding complexity only when the work demanded it
- Persistent project rules instead of re-explaining
Failed:
- Same-model code review. Confirmation bias, exactly like a developer testing their own code. This one became a whole post: Reviewer design: catching what the architect tier misses.
- Casual prompts ("continue", "keep going"). My notes say I lost most of the working context immediately.
- Architectural decisions without domain input. Over-engineered every time.
- Too many agents up front. Overhead, no value.
- Long, uninterrupted sessions. Coherence degraded with length.
- Expecting memory across repositories. It isn't there unless you build it.
- Trusting AI-generated tests without review. The tests adapted to the code, not to the requirements. That one still shapes how I build the test suite today.
How the human role changed
What disappeared: writing code by hand, routine debugging, repetitive refactoring, boilerplate, basic docs.
What appeared: designing the orchestration, systematic prompting, managing context, deciding which agent does what, setting architectural direction, judging quality, and applying the domain knowledge nobody else in the "team" has.
The split I settled on: the human owns the vision, the domain, user empathy, the orchestration and the ship-or-iterate decision. The AI owns implementation, framework expertise, pattern consistency, test volume and documentation.
What transfers if you're a tester
My testing background turned out to be the most useful thing I brought. The principles carry over almost unchanged:
- Separation of concerns. Different agents, different responsibilities.
- Independent review. Never let the same model mark its own homework.
- Systematic checks. Automated, but human-verified.
- Requirements-driven testing. AI tests the code it wrote, not what you needed.
- Risk-based prioritisation. Spend AI effort where failure costs most.
And there is new work for us: testing the orchestration itself for consistency, validating prompts, detecting context drift, comparing models task by task. We already think in edge cases, question assumptions, and know that "it works" is not the same as "it's right". That mindset is an unfair advantage here.
Where this went next: the quality side grew into ~70,000 tests and a verification repo the implementation can't touch. See Quality at machine speed.
Talk materials
The first public version of this story was a Testing United talk, "Is AI the future of Software Delivery: Can it replace a Feature team?" [MIKE: confirm year, 2024 or 2025, and whether to link the slides PDF here.]
The deck ran to 37 slides in three acts: the problem (an archer's dilemma, the constraints), the journey (four phases, what worked and failed in each), and the insights (team structure, the human role, economics). My favourite opening line from it still holds: twenty years in testing, zero iOS experience, and a production app at the end of it.
[MIKE: attach or remove — slides PDF, agent prompt templates, rules-file example, phase timeline graphic.]
Final thought, then and now
In November 2025 I ended these notes by saying AI didn't replace my expertise, it amplified it. Eleven months, an App Store launch and a book chapter later, I'd sharpen that: it amplified whatever I brought, including my blind spots. Which is why most of what I've written since is about building the checks that catch them.
If you're running your own version of this experiment, where did your workflow hit its first wall?
Last updated: November 2025 Part of the AI Orchestration Journey series