A local-only Monte Carlo retirement simulator, built to answer one real question with numbers instead of fear: given my actual accounts, spending, and taxes, what's my probability of never running out of money - and what's the most I can safely spend per year?
Spending is modeled as what you actually live on, after tax - the engine grosses up each withdrawal to cover the tax owed on it, so the number you enter is the number you spend, not a pre-tax abstraction. Guardrails cut spending automatically when a plan drifts off course, the same way a real financial plan would, never letting "success" mean anything less than covering essential, non-discretionary spending first.
Priorities, in order: correct, pretty, fast - deliberately in that order, because the entire point of the tool collapses if the math is wrong. Every tax and financial function's docstring states the rule and its source; a deterministic single-path engine exists purely as a readable reference the vectorized Monte Carlo engine is tested against, numerically, for exact equivalence.
Nothing selected yet — click a box in the diagram above.
Here's the concrete problem the MCP server solves: the advisor and the engine are two separate running programs, on two separate ports - the advisor can't just import the engine and call its function directly, the way you could inside one Python program. It has to reach across that process boundary somehow. The MCP server is that bridge, and all it does is describe one side to the other - it adds no math, no new logic, nothing that could get the numbers wrong.
1. The real function already exists in backend/engine/ - ordinary Python,
nothing AI-aware about it:
def run_baseline(accounts, spending, assumptions):
... # the actual 5,000-path Monte Carlo
return paths # a plain Python return value
2. The MCP server publishes a description of that one function - its
name, and the shape of its inputs and outputs - as a "tool":
tool name: run_baseline
takes: accounts, spending, assumptions
returns: 5,000 simulated paths
3. The advisor calls the tool by name over the network instead of
importing the function:
result = mcp_client.call("run_baseline", {accounts, spending, assumptions})
"result" is exactly what run_baseline(...) would have returned if
the advisor could have called it directly - the MCP server just
relayed the call and the answer across the process boundary.
So "wrapping as a callable tool" is step 2 alone: writing down a function's name and its input/output shape so something outside the process can ask for it by name. The function in step 1 is unchanged either way - it's tested and used the same whether it's called directly (the deterministic reference engine does this) or reached through the MCP server (the advisor does this).
The math is never AI. The engine is pure NumPy and stdlib - deterministic, auditable line by line. The advisor is a separate layer on top: a RAG-grounded chat that reads the engine's already-computed output and, only with explicit consent, can propose one specific change back into the plan.
A retirement plan is genuinely sensitive financial data. The simplest way to never leak it from a server is to never let it reach one.
There is no server-side database. Every plan lives in the browser's own IndexedDB and the backend forgets it the instant it responds. The one thing that is server-side is the advisor's own reference library - synthetic retirement-planning material, never a real household's numbers.
Everything above covers what Runway is. From here on is how it actually works - the real architecture, the tax-and-Monte-Carlo engine, where household data actually lives, a real bug in the AI advisor with its own regression test, a stale-data incident and the design that fixed it, and the test suite that backs all of it. Every claim below is sourced from Runway's own commit history and source code, not written from memory.
React, Vite, and TypeScript on the frontend; FastAPI on the backend; nothing else in between. The engine itself - backend/engine/ - is a hard architectural boundary: it imports only NumPy and the Python standard library, no FastAPI, no web framework, so the math can never accidentally depend on how it's served. Every plan lives in the browser's own IndexedDB; the backend holds no household data of its own, no server-side database at all. A household's real profile, when one exists, lives in a single local JSON file, git-ignored and never sent anywhere - the same discipline the codebase's own rules state directly: never commit it, never send it to an external service, never paste its contents anywhere.
An optional AI advisor runs as its own separate FastAPI service (service.py), holding an in-memory job queue - no database, since a job only needs to live as long as one request is waiting on it. Generation never calls out to FOMX.ai's Spark directly: a worker running on the Spark polls that queue over Tailscale, claims the job, calls Ollama locally, and posts the result back - the same direction-reversed pattern every app on this platform uses, so the Spark itself never has to accept an inbound connection. It is deliberately not required for the core tool to work, and it never computes a number - it can only narrate results the deterministic engine already produced. That last guarantee isn't just policy: the advisor reaches the engine through a third process, an MCP (Model Context Protocol) server (mcp_server.py), whose only job is exposing the engine's existing functions as typed, callable tools. The advisor has no other path to the math - it literally cannot import backend/engine/ directly, so a number reaching the model has to have come from the real engine, through a real tool call, with a real schema.
Priorities, stated in the codebase itself, in order: correct, pretty, fast - because the entire point of the tool collapses if the math is wrong. A deterministic, single-path engine is the readable reference implementation; a vectorized Monte Carlo engine (5,000 simulated paths, roughly 41 years each, in under two seconds) is tested to numerically match it. Every tax and financial function's docstring states the rule and its source - not just what the code does, but which statute or IRS table it's implementing.
Spending is modeled as what a household actually lives on, after tax - the engine grosses up each withdrawal to cover the tax owed on it, so the number a person enters is the number they actually spend, not a pre-tax abstraction. Tax brackets and deduction amounts grow with each simulated path's own realized inflation; Social Security claiming age rescales the benefit per SSA rules. Guardrails cut spending automatically when a plan drifts off course, the same way a real financial plan would - success is never allowed to mean anything less than covering essential spending first.
No server-side database, anywhere. The backend's config_store.py loads and saves a household profile to a local JSON file - one per named profile, so the same install can hold real data alongside placeholder demo data, switchable from the UI. Saves are atomic: written to a temporary file and then swapped in, so an interrupted save can never corrupt a profile. In the deployed, multi-user version, plan state moves client-side entirely, into the browser's own IndexedDB - the backend never holds it at all. Real household data, when it's in use, never leaves the machine and is never committed to source control.
Someone asked the advisor a simple question: "what if I get an inheritance in 3 years?" The AI misread "3 years from now" as the literal age 3 - decades before the person's real age - instead of doing the math to convert it. Because no year in the simulation ever hits age 3, the tool quietly compared the plan to itself instead of testing the change, and confidently reported that an $800,000 inheritance would make zero difference. The math ran without error. The answer was simply wrong, and it looked exactly like a right one. (Verified against the actual fix in the codebase, commit 87a3d9e.)
This is the failure that matters most for a tool built on one promise: the AI explains the numbers, it never invents them. A crash is obvious and gets noticed immediately. A confident, wrong answer isn't - which is exactly why it's now caught automatically, every time, before it ever reaches a user.
A separate, already-documented incident: the dashboard once showed two numbers side by side that contradicted each other - a success rate in the low 70s next to a "safe to spend" figure that was actually higher than what the plan called for. Each number had been correct at some point, just not at the same time. The cause: turning Social Security and inheritance income off should have triggered a fresh calculation, but the app's caching logic didn't notice anything had changed, so it kept showing the old, now-outdated result next to the new one.
The fix: every result the system produces is now tagged with a unique fingerprint, and the dashboard is built to refuse to display two numbers together unless they carry the same one. That rule is now written into the team's own engineering standards and backed by a permanent automated test, so this kind of silent mismatch can't slip through again unnoticed.
324 backend tests and 73 frontend tests actually run and pass on every single push to the codebase - not written and forgotten, but continuously enforced. A production build and a separate type-check act as two more automatic gates a change has to clear. The most important tests compare two independent versions of the same math against each other - a fast production engine and a slower, simpler reference version built purely to double-check it - and fail loudly if they ever disagree, down to a tight tolerance. The team rule behind them is simple: change a calculation, update both versions, or the tests catch the mismatch automatically.
Beyond that, testing is organized into tiers by what each one protects: one tier locks in exact, known-correct results for a fixed reference case, so any change that would quietly shift a number the business relies on fails immediately; another deliberately throws malformed or garbage input at the system to confirm it never breaks; a third actively tries to trick the AI advisor into writing data it shouldn't be able to write, to confirm that safeguard actually holds. What's deliberately not done - and why - is written down too: a few advanced testing techniques were skipped as low-value at this project's current size, a documented judgment call rather than a gap nobody noticed.