Engineering Platform › Shared module · See it on the architecture map

AI Answers (RAG)

RAG stands for Retrieval-Augmented Generation. In plain terms: instead of just asking an AI a question and hoping it remembers the right answer, this system first goes and finds the actual, relevant paragraphs from real documents, hands the AI only those paragraphs as evidence, and then double-checks that whatever the AI says it "found," it actually found. Built for cases where the AI answering the question is a small, cheaper, self-hosted model - the kind that's more prone to making things up than a large one. Runway's financial-planning advisor is its real, live user.

Why it matters
An AI that names a source it never read looks trustworthy and is lying. In a financial tool that is worse than no answer at all - so the answer has to be grounded in real documents and the citations have to be checked, not trusted.
What I designed
Ingestion, chunking, embedding, retrieval and citation enforcement as one standalone package; pointing it at a different document set is the only thing that changes between applications.
Used by
Runway's advisor.
Finds text by meaning
not by matching keywords - it understands that "monthly cost" and "premium" are related ideas
Citations are checked, not trusted
if the AI names a source that was never actually used, that source is deleted before the answer is shown
1 real caller
Runway's advisor uses this in production

Why this exists: AI models sometimes make up their sources

Runway's advisor is told: "answer the question, then tell the user which document you got that answer from." Big, expensive AI models (like the ones behind ChatGPT or Claude) are usually honest about this. Smaller, cheaper models - the kind that can run on Derek's own computer instead of paying per question - often aren't. They'll confidently name a document that sounds exactly right and was never actually shown to them. The package's own test suite catches a real example of this: a model inventing the source name deductible_definition.txt out of thin air. A made-up citation is worse than no citation at all - it looks trustworthy, and it's a lie.

The fix isn't "ask the model more nicely." It's checking its homework. The system already knows, with total certainty, which real documents it actually handed to the model, because it fetched them itself. So after the model finishes its answer, a separate step called enforce_citations() compares what the model claims it used against what it was actually given - and throws out any citation that doesn't match. What's left is rebuilt into one clean, accurate "sources" line. Just as important: it never lists every document that was available, only the ones actually used - an earlier version tried that and it backfired, because "I don't have an answer for you" would still show five unrelated document names, making a non-answer look like evidence. No source listed is more honest than a misleading one.

How it actually finds the right paragraph

Think of it in four steps. First, every document gets cut up into small chunks - a paragraph or so at a time, small enough to be a focused answer to one question. Second, each chunk is converted into a list of numbers that captures what it means, not just the words in it - two chunks about the same idea end up with similar numbers, even if they don't share a single word. Third, when a user asks a question, that question gets turned into the same kind of number list. Fourth, the system compares the question's numbers against every chunk's numbers and pulls out whichever chunks are the closest match - those become the "evidence" handed to the AI, and the citation check above runs on whatever it does with them.

That comparison step uses something called cosine similarity - which sounds technical but is a simple idea: it measures how closely two things point in the same "direction" of meaning, the way you might say two arrows are pointing roughly the same way even if one is longer than the other. That matters here because a short question and a long paragraph shouldn't be penalized just for being different lengths - cosine similarity only cares whether they're "about" the same thing, which is exactly what you want when matching a question to an answer.

Generating the actual answer text goes through a swappable connector inside the package - by default a small local model running on FOMX.ai's GPU machine (the "FOMX.ai Spark"), or Claude itself, reached over the internet, when Runway chooses to use it instead. That swap point is exactly the kind of thing a future shared module (model_kit, planned but not yet built) would handle for every app at once, instead of each app wiring it up separately.