This is a solo-developer project, but no change reaches players by going straight from an idea to production. Every change moves through the same 11-stage pipeline below - I open it and I close it, with an Architect, a Developer, a Tech Writer, and an independent reviewer running in between, each one an AI session under a written contract instead of tenure.
Eleven stages, one fixed order. I hold three of them myself - the ones a role contract can't substitute for; an AI role session holds the rest, each one closed and reopened blank for the next item.
/projectplan) - starts as a conversation, no file yet./architect) - the conversation ends here; acceptance criteria, owned files, and the invariant the change must preserve, committed to specs/./dispatch <ID>) - reads the spec's own Owner: field and picks Developer or Tech Writer itself./build <ID>) - exactly the spec, in the files it names; anything the spec is silent on stops and goes back, not improvised.review_with_qcoder.py) - a separate model call per changed file, seeing only the file's current content and the changed lines, never the reasoning behind the change./review <ID>) - re-verifies the report's own claims against the code directly, records a verdict./build <ID> on the linked docs spec) - same commit-and-report discipline as a Developer build.git push) - secret scan, tests, lint, security-lint, dependency audit; the deploy job won't run unless every one of them succeeded./test <ID>).promote.sh sit) - a full, isolated copy of the app, so testing never collides with active development./accept <ID>) → production.This runs like a real software team - a product owner, an architect, developers, a tech writer, a reviewer, QA - with most of those seats filled by AI, working under a written contract instead of tenure.
| On a human team | Here | Filled by | Never |
|---|---|---|---|
| Product owner · BA · project manager | The Principal | A person | Nothing - final, non-delegable authority |
| Backlog sequencing · status reporting | The PM Session | AI | Dispatch work, write a spec, or record a verdict |
| Solution architect · tech lead | The Architect | AI | Write or delegate product code |
| Developers | Developer(s) | AI, one or several | Redesign scope - kicks back instead |
| Technical writer | The Tech Writer | AI | Edit anything outside its own docs |
| Code reviewer | Qcoder, then the Architect | Tools + AI | Substitute for the Architect's own review |
| QA · user acceptance testing | Business testing in SIT | A person | — |
Segregation of duties is the load-bearing rule: the session that defines the work, the session that does it, and the person who accepts it are never the same one. It matters more here than on a human team - a person who's unsure tends to hesitate; an AI that's wrong sounds exactly like an AI that's right.
A session doesn't know what it is until it reads a file. Nobody opens a "developer window" - a session opens blank, is told which spec to build, reads the dispatch note naming its role, and works under that role's contract for the rest of the session.
Three refusals are hard, every time:
Every open item is a file with a header - state:, owner:, blocked_by:, files:. Nobody reads dozens of those by hand; one command computes the state of everything from the headers themselves, never from memory.
$ python3 tools/board.py --next
NEXT: /review T36
2 verdicts left: T36, T91
then 10 startable: T84, T89, T95, T56, T67, T68, T71, T93, T77, T94 (22.25h)
NEEDS YOU: T7, T9
25 open, 41.2h total
A real morning's output. Five lines answer the only question that matters: what's next, what's waiting on a verdict, what's safe to run at the same time, and - read first - what needs a decision only a human can make.
Four rules decide what's next, applied in order, stopping at the first that fires:
Sequencing itself is three commands that require judgment about order, never about the code: /projectplan shows state, /dispatch writes the briefing and picks Developer-or-Tech-Writer from the spec's own owner field, /build is what a fresh session types to pick it up. And running several sessions at once adds a second failure mode this same file-ownership check closes: a spec declares exactly which files it owns before anything is dispatched, and an item whose files overlap with whatever's already running does not get handed out, even if nothing else about it is wrong.
Four steps, most of the time run smoothly: build (with every machine check), hand off, ask for review, get a verdict. The rule for when it isn't smooth is the same rule that governs everything else here - a Developer never resolves something outside their own spec. They document what they found and hand off; the Architect resolves it or writes the follow-up as its own item, and the verdict says so explicitly - conforms-with-deviations, or kicked-back with exactly what's owed. Never a silent "looked fine to me."
A real piece of work: T20, closing a gap in a settings leak-guard's own test coverage - the formatter was already unit-tested, but nothing yet proved the guard's pytest.fail actually fires on real drift.
/dispatch T20conftest.py read-only), what must not break (11 existing tests), and what's explicitly not in scope. Five lines; the two that matter are "must not break" and "not doing" - where a wrong instruction is visible before any code exists./build T20pytest.fail silently mutated to pass) makes one specific test fail. All fourteen tests pass on the real file./review T20conforms. Four commands and one ID typed by a human. Nothing copied between windows, nothing decided without sign-off, and every step left a committed file behind - so if any session had died mid-way, the next one could pick up exactly where it stopped.Eleven independent checks stand between a session's first keystroke and anything reaching a server. Listing what each catches and stopping there would be a brochure - the column that matters is the one on the right.
| Safeguard | Catches | Does not catch |
|---|---|---|
| Architect review | Conformance to spec, proven by mutation | Whether the spec was right |
| Qcoder LLM review | Defects in .py | Templates, workflows, shell, config - no review at all |
pytest -q on every push | The full test suite | Nothing end-to-end |
ruff check . | Lint and style | — |
bandit | Insecure Python patterns | Excludes tests/ |
pip-audit | Vulnerable dependencies | — |
scan_secrets.py + pre-commit hook | Committed secrets | Values that aren't quoted string literals |
Both deploy jobs needs: all five gates | Anything deploying un-green | — |
promote.sh green check | Non-green shas - refuses on "no checks at all" and "still running" too | — |
environment: production reviewer | An unreviewed production release | It's still one person approving their own release |
| Prod fast-forwards to the SIT tip only | New code entering prod without passing SIT | — |
Two of these are worth calling out without hedging: the green check treats "no checks have run at all" as a refusal too - exactly the case a naive "nothing failed" test gets wrong - and both deploy jobs require all five CI gates before they'll run, so SIT is protected by exactly the same bar as production, not a lighter one.
Four incidents and the rules they produced are already told on the main engineering page - a fifth belongs here, not yet told there:
A doc page can drift from the code it describes just like code can drift from its spec - so the same discipline applies to both. When a build changes user-facing behavior, the Architect writes the docs follow-up in the same review pass, not as a note for later; the Tech Writer's own draft then runs a self-review loop (writer → skeptical student → professor → writer again) before it ships, because confidently-wrong prose passes every automated gate built to catch confidently-wrong code.
SIT - the isolated rehearsal stack I test in myself before anything reaches players - has been verified the same way, on real dates, not assumed correct because it deployed:
200 - checked directly, not inferred from config.