I don't build one application at a time. Underneath Tourney, Runway and GetSeen is one engineering platform: the same CI/CD discipline, the same production host, the same GPU inference setup, and reusable code built once and used everywhere. The applications are the proof; the platform is the point.
00 Why it matters, and what I designed
Why it matters. An application that works once is a demo. An application that is built, tested, deployed, secured and operated the same way as its siblings - by a process that refuses to skip a step - is a system someone else could run. That is the standard I held enterprise platforms to at TD Bank, and it is the standard here.
What I designed. One control plane, one pipeline, and a set of shared modules that turn repeated problems into reusable components:
Control plane
Control Tower - status, one governed path to production, fail-closed gate, append-only audit log, AIOps investigation of failed releases.
New apps
A guided + New app flow checks the setup, then builds the app with Dev and SIT already wired into the same pipeline. Prod is never part of setup - it stays a human click.
Pipeline
Five gates (secrets scan, tests, lint, security-lint, dependency audit); DEV → SIT automatically on green, SIT → PROD by a deliberate human trigger. Section 01 below.
Source control
Nine private repositories on one account: one per app, one for the control plane, one per shared module, one for the shared CI. Branches map to environments. How it is organized.
Shared kits
Each chore every app needs (deploys, checks, packaging, web address, health, backups) is written once and used by every app, new ones from day one. Why, and what each does.
Backups
A nightly copy of production data, proven by restoring it, kept in a private encrypted bucket, with the last-backup time shown in Control Tower. Designed and approved, being built. How it works.
Infrastructure
One AWS host behind one nginx for every live application; TLS, Docker networking and deploy identity shared on purpose. Section 02.
AI backends
Self-hosted open-weight models on an NVIDIA DGX Spark for volume, a frontier model where quality earns its cost - every workload run through both first. Section 03.
The control plane, running. Click to see what each part does.
The rest of this page is how it works. The module pages are the technical detail.
01 I run real CI/CD, not "it works on my machine"
I put every deploy through the same five gates: a secrets scan, the full automated test suite, lint, security-lint, and a dependency audit - all before anything reaches a live environment. I wired it so a commit to main auto-promotes to system integration the instant those gates go green, and I keep production promotion a deliberate, separate step - never automatic, never skipped.
Stated plainly, not glossed over: Tourney and Runway now run the same four shared gates, gate for gate, through one reusable pipeline (fomx-ci) - each still keeps its own local test suite as the fifth, since the two apps don't share a test runner. I closed the last gap (lint and security-lint on Runway) by diffing the two workflows line by line rather than assuming they matched, fixing one real bug that surfaced along the way (a missing environment-forwarding flag that would have silently broken Runway's very first production deploy).
Promotion itself is one deliberate action, not a form to fill out: Actions tab → "Promote to prod" → Run workflow (or gh workflow run promote-to-prod.yml), defaulting to SIT's current tip. That single trigger verifies every gate is green, fast-forwards the prod branch to that exact commit, kicks off the real deploy, and confirms the live app answers healthy - I did this for real on Runway's first production promotion, and hit (then fixed) a genuine gap along the way: a GitHub-token push doesn't cascade into another workflow run, so the promotion step has to explicitly trigger the deploy rather than assume the branch move alone will do it - the same fix Tourney's own pipeline already had, discovered the same way, on its first promotion too.
02 I run production infrastructure, not just code
I host multiple live applications on one AWS EC2 box, side by side, behind one nginx - I didn't stand up a second server for the second app; I added a container and one nginx config file, and nothing already running had to change. I did this myself, live, on the production box: added a new app alongside the existing one, got it its own TLS certificate, and verified the existing app never went down.
I issue and renew TLS certificates myself via Let's Encrypt/certbot, expanding a certificate to a new hostname rather than replacing it.
I join a new app to the existing Docker network as an internal service - it never gets its own exposed port, only reachable through nginx.
I diagnosed and fixed a real production incident live - a stale Docker bind-mount that silently kept serving an old nginx config - by isolating it methodically rather than guessing.
I share deploy credentials on purpose, not by default - the same host and the same restricted SSH key log every app's CI/CD pipeline into this box, and the same GitHub token (already scoped to every repo I own) checks out each app's own code once it's there. Adding a third app here, I caught myself about to mint yet another single-repo deploy key out of habit - checked first, confirmed the existing token already covered it, and used that instead. Fewer secrets to rotate and audit, same principle as reusing the box itself.
03 I evaluate AI cost and quality, I don't just call an API
I operate FOMX.ai's dedicated NVIDIA DGX Spark - the company's own physical GPU hardware - serving self-hosted, open-weight models through Ollama, and I put every AI workload through both that and a frontier model (Claude) before deciding which one runs it in production. That's a real cost/quality tradeoff I make deliberately, not a default I never revisited.
I use frontier models where quality matters most and the volume is low enough that the per-token cost is worth it.
I use self-hosted models on FOMX.ai's hardware for high-volume or cost-sensitive work, where a step down in quality is an acceptable trade for zero marginal cost.
I designed the connection so the GPU box never accepts an inbound connection - a worker I built polls a job queue and pulls work on its own schedule, so the model host is never directly reachable from the internet.
04 I build reusable platform code, not one-off scripts
When I hit the same problem in a second application, I don't solve it twice - I extract what I built the first time into its own package and reuse it.
AI Tool Loop (agentkit)
I generalized a job-queue, tool-using model loop I originally built for one app into a portable package any app can drop in - the queue, the tool registry, and the security boundary between an app and a worker it doesn't fully trust.
I built document ingestion, chunking, embedding, retrieval, and citation enforcement once, as a standalone package - pointing it at a different document set is the only thing that changes between applications.
I extracted the client that talks to FOMX.ai's self-hosted GPU box - the HTTP polling loop, the claim logic, the usage tracking - so every app that needs the FOMX.ai Spark reuses the same battle-tested code instead of its own copy.
Not everything generalizes cleanly, and I don't pretend it does. When I looked at reusing Tourney's login for a second app, the two apps turned out to be built on different foundations - Tourney is a Flask app, with Flask's own session cookie and password-hashing conventions baked in; the second app is a lighter, framework-free HTTP server with none of that machinery. Copying Tourney's login code directly wouldn't have worked. What I actually reused were the two underlying, framework-independent primitives Flask's own login is built from - password hashing (werkzeug) and signed session tokens (itsdangerous) - and wired each app's very different server up to them the same way. Same security primitives, two different integrations, because pretending the frameworks matched would have been the wrong kind of reuse. That extracted package is Auth Kit →.
05 I build for maintainability, including my own future self
Every app I run has exactly one supported way to start it in development - never a bare command I have to remember the right flags for.
I deliberately turned off live-reload where I'd seen it cause real problems: on Windows, I found a dev server's reloader orphaning a child process and silently wedging its own port - so I made a real restart the rule there, not a hot reload I couldn't trust.
On another app, I know its Python modules only load once at startup - editing code on disk changes nothing a running server serves until it's actually restarted - so the launch script stops old processes first, every time, rather than relying on anyone remembering to.
The pattern I apply everywhere: one named script is the only way I start dev, specifically so "did you actually restart it" is a problem I solved once, not a debugging session I repeat.
06 I pick the storage architecture the requirement actually needs
Every promotion to SIT or production across my apps goes through one tool, Control Tower, and every one of those actions gets recorded - who approved it, when, which commit, which environment, whether it actually succeeded. The interesting decision wasn't the UI. It was what to store that record in.
Database, or a file? I chose an append-only event log - one immutable JSON record per line, never edited, never deleted - over a database, and I can defend exactly why. This is event sourcing: the log itself is the source of truth, not a side effect of some other table. A mistake gets corrected by appending a new record, never by rewriting history - the same discipline real audit-log systems use specifically because an append-only structure makes tampering visible by construction, not just by policy. A database would have given me transactions and indexed queries I don't have volume to need yet; the file gives me a property a database doesn't hand you for free: the storage format itself cannot represent an edit. I designed the record schema so a move to a real database later is a storage-layer swap, not a rewrite - right-sized to the actual requirement, not the default choice.