Process

This is a solo-developer project, but no change reaches players by going straight from an idea to production. Every change moves through the same 11-stage pipeline below - I open it and I close it, with an Architect, a Developer, a Tech Writer, and an independent reviewer running in between, each one an AI session under a written contract instead of tenure.

11-stage pipeline
a fixed order, every change, no side branches
6 roles, 1 human
segregation of duties - the same rule real software teams run on
11 safeguards
independent checks between a keystroke and a live server
4 prioritization rules
what runs next, computed from spec headers - never from memory

The pipeline, end to end

Eleven stages, one fixed order. I hold three of them myself - the ones a role contract can't substitute for; an AI role session holds the rest, each one closed and reopened blank for the next item.

  1. Administrator brings the ask (/projectplan) - starts as a conversation, no file yet.
  2. Architect writes the spec (/architect) - the conversation ends here; acceptance criteria, owned files, and the invariant the change must preserve, committed to specs/.
  3. Administrator dispatches to the builder (/dispatch <ID>) - reads the spec's own Owner: field and picks Developer or Tech Writer itself.
  4. Developer builds, with tests in the same commit (/build <ID>) - exactly the spec, in the files it names; anything the spec is silent on stops and goes back, not improvised.
  5. Independent AI review (review_with_qcoder.py) - a separate model call per changed file, seeing only the file's current content and the changed lines, never the reasoning behind the change.
  6. Architect reviews the actual diff (/review <ID>) - re-verifies the report's own claims against the code directly, records a verdict.
  7. Tech Writer updates the docs the diff touched (/build <ID> on the linked docs spec) - same commit-and-report discipline as a Developer build.
  8. CI/CD - mechanical, every push (git push) - secret scan, tests, lint, security-lint, dependency audit; the deploy job won't run unless every one of them succeeded.
  9. Administrator checks what to test and read (/test <ID>).
  10. Administrator tests it directly, in SIT (promote.sh sit) - a full, isolated copy of the app, so testing never collides with active development.
  11. Administrator approves it for production (/accept <ID>) → production.
Kick-backs run backward, on purpose. A later stage can send work back to the Architect - an ambiguous spec, a diff that doesn't match what was promised, a defect found in SIT - and every one of those lands on a spec-level question, not a quick patch. A defect found in SIT is exactly as serious as one caught in review; neither gets waved through because it's "probably fine."

One team, mapped onto AI

This runs like a real software team - a product owner, an architect, developers, a tech writer, a reviewer, QA - with most of those seats filled by AI, working under a written contract instead of tenure.

On a human teamHereFilled byNever
Product owner · BA · project managerThe PrincipalA personNothing - final, non-delegable authority
Backlog sequencing · status reportingThe PM SessionAIDispatch work, write a spec, or record a verdict
Solution architect · tech leadThe ArchitectAIWrite or delegate product code
DevelopersDeveloper(s)AI, one or severalRedesign scope - kicks back instead
Technical writerThe Tech WriterAIEdit anything outside its own docs
Code reviewerQcoder, then the ArchitectTools + AISubstitute for the Architect's own review
QA · user acceptance testingBusiness testing in SITA person—

Segregation of duties is the load-bearing rule: the session that defines the work, the session that does it, and the person who accepts it are never the same one. It matters more here than on a human team - a person who's unsure tends to hesitate; an AI that's wrong sounds exactly like an AI that's right.

The Principal (human) PM Session Architect Developer Tech Writer Code Review (Qcoder + Architect)

The role lock

A session doesn't know what it is until it reads a file. Nobody opens a "developer window" - a session opens blank, is told which spec to build, reads the dispatch note naming its role, and works under that role's contract for the rest of the session.

Three refusals are hard, every time:

The contract is the mechanism. The role lock works through discipline, not a runtime permission system - every session reads the dispatch note, adopts the named role, and holds to its contract for the rest of the build.

Working the queue

Every open item is a file with a header - state:, owner:, blocked_by:, files:. Nobody reads dozens of those by hand; one command computes the state of everything from the headers themselves, never from memory.

$ python3 tools/board.py --next
NEXT:  /review T36
2 verdicts left: T36, T91
then 10 startable: T84, T89, T95, T56, T67, T68, T71, T93, T77, T94  (22.25h)
NEEDS YOU: T7, T9
25 open, 41.2h total

A real morning's output. Five lines answer the only question that matters: what's next, what's waiting on a verdict, what's safe to run at the same time, and - read first - what needs a decision only a human can make.

Four rules decide what's next, applied in order, stopping at the first that fires:

  1. An unreviewed build outranks a new one - a commit sitting without a verdict often blocks other work behind it.
  2. Whatever unblocks the most work goes next - an item three others are waiting on outweighs a standalone item twice its size.
  3. A kicked-back item before a fresh one - it's already mostly done and already verified once.
  4. Only then, by value - the board's own "startable" list is deliberately unordered beyond that; judging which is worth doing next is a decision a person makes, not a script.
Why this exists: it isn't hypothetical. Before the board did, one day's work got hand-scheduled eight separate times and was wrong five of those times - dispatched onto an item with no briefing, a paused item, a held item, marked "waiting" on a phantom blocker that shared no files at all, and once, the same item handed to two sessions at the same time on the same filesystem. Every one of those was information already sitting in the repo; a session just had to read five files to find it, and didn't.

Sequencing itself is three commands that require judgment about order, never about the code: /projectplan shows state, /dispatch writes the briefing and picks Developer-or-Tech-Writer from the spec's own owner field, /build is what a fresh session types to pick it up. And running several sessions at once adds a second failure mode this same file-ownership check closes: a spec declares exactly which files it owns before anything is dispatched, and an item whose files overlap with whatever's already running does not get handed out, even if nothing else about it is wrong.

The build loop, and what happens when something's wrong

Four steps, most of the time run smoothly: build (with every machine check), hand off, ask for review, get a verdict. The rule for when it isn't smooth is the same rule that governs everything else here - a Developer never resolves something outside their own spec. They document what they found and hand off; the Architect resolves it or writes the follow-up as its own item, and the verdict says so explicitly - conforms-with-deviations, or kicked-back with exactly what's owed. Never a silent "looked fine to me."

One item, start to finish - every keystroke

A real piece of work: T20, closing a gap in a settings leak-guard's own test coverage - the formatter was already unit-tested, but nothing yet proved the guard's pytest.fail actually fires on real drift.

/dispatch T20
The Architect writes the full briefing to a file: what's building, which one file to touch (tests only, conftest.py read-only), what must not break (11 existing tests), and what's explicitly not in scope. Five lines; the two that matter are "must not break" and "not doing" - where a wrong instruction is visible before any code exists.
/build T20
Six characters. A fresh session reads the briefing, takes the Developer role because the briefing says so, writes three new tests, and proves each one by mutation against a scratch copy - the exact disarm that motivated the spec (pytest.fail silently mutated to pass) makes one specific test fail. All fourteen tests pass on the real file.
/review T20
The Architect reads the write-up, then ignores it and checks for itself - re-runs all fourteen tests, reproduces the mutants, and looks at one thing nobody was asked to look at. That habit paid off: the guard's own registration (whether pytest actually wires the fixture in) was untested by either this build or the one before it. Raised as a new item, not fixed on the spot - deciding scope is Architect work.
Result
Verdict: conforms. Four commands and one ID typed by a human. Nothing copied between windows, nothing decided without sign-off, and every step left a committed file behind - so if any session had died mid-way, the next one could pick up exactly where it stopped.

What each safeguard catches - and what it doesn't

Eleven independent checks stand between a session's first keystroke and anything reaching a server. Listing what each catches and stopping there would be a brochure - the column that matters is the one on the right.

SafeguardCatchesDoes not catch
Architect reviewConformance to spec, proven by mutationWhether the spec was right
Qcoder LLM reviewDefects in .pyTemplates, workflows, shell, config - no review at all
pytest -q on every pushThe full test suiteNothing end-to-end
ruff check .Lint and style—
banditInsecure Python patternsExcludes tests/
pip-auditVulnerable dependencies—
scan_secrets.py + pre-commit hookCommitted secretsValues that aren't quoted string literals
Both deploy jobs needs: all five gatesAnything deploying un-green—
promote.sh green checkNon-green shas - refuses on "no checks at all" and "still running" too—
environment: production reviewerAn unreviewed production releaseIt's still one person approving their own release
Prod fast-forwards to the SIT tip onlyNew code entering prod without passing SIT—

Two of these are worth calling out without hedging: the green check treats "no checks have run at all" as a refusal too - exactly the case a naive "nothing failed" test gets wrong - and both deploy jobs require all five CI gates before they'll run, so SIT is protected by exactly the same bar as production, not a lighter one.

A fifth incident

Four incidents and the rules they produced are already told on the main engineering page - a fifth belongs here, not yet told there:

What happened
Two Architect sessions, reading the same stale view of the spec folder, independently stamped the same two ID numbers onto four different specs. A dispatch command can't survive a duplicated ID - it dispatches whichever file it reads first.
Rule now
An ID is allocated in one file and nowhere else. The claim is committed alone, before the spec itself is written - it doesn't prevent a race, it makes one cheap: one conflict in one small file instead of a contradiction spread across three places.

Verified, not assumed

A doc page can drift from the code it describes just like code can drift from its spec - so the same discipline applies to both. When a build changes user-facing behavior, the Architect writes the docs follow-up in the same review pass, not as a note for later; the Tech Writer's own draft then runs a self-review loop (writer → skeptical student → professor → writer again) before it ships, because confidently-wrong prose passes every automated gate built to catch confidently-wrong code.

SIT - the isolated rehearsal stack I test in myself before anything reaches players - has been verified the same way, on real dates, not assumed correct because it deployed: