Skip to content

Architecture

Design for the full system. Everything in the diagram below is built: the safety gate, the confidence floor, retrieval, the LLM step, validation, the handover email, the HTTP entry point, a React chat in front of it, and the upload screen a mentor updates the knowledge base from. Still to come: the Disciple adapter, which is the second surface. Task-by-task status is in index.md; this page is about why the pieces are shaped as they are.

One branch in this diagram is not what the original design called for. It predicted two confidence thresholds — one for off-topic and one for "on-topic but the knowledge base does not cover it". Measurement killed the second, and the reasoning is under "One threshold, not two" below.

Request flow

Every branch has a defined destination. That is the point of the diagram: no path is left undefined, so there is no state in which the bot simply does nothing.

Surfaces (React chat | Disciple adapter, to come)
        |
   FastAPI /ask            <- single entry point, both surfaces
        |
   Orchestrator (ask.py)   <- owns the MLflow trace and response assembly
        |
   Safety gate ------------> sensitive: office number + email team
        |            ------> crisis: fixed 24/7 resource response, stop
        |
   Confidence probe -------> cosine < 0.58: canned scope reply, no email
        |                    (off-topic questions need nobody to follow up)
        |
   Retrieve top 4           <- cross-encoder rerank
        |
   DSPy + LLM -------------> no lesson fits: "can't help" scope reply, no email
        |                    (this is what catches "on-topic, not covered")
        |
   Validate ---------------> fails: email team, log as bug
        |
   Reply (back through orchestrator)

The email is sent after the reply goes out, as a background task, so a slow mail server costs the member nothing. Measured latencies through the whole path: 0-1 ms when the gate fires, ~30 ms for an off-topic reply, ~2.4 s when the model runs — of which ~1.9 s is Groq. Everything that is not the LLM call is close to free.

Ingest is a separate path

Ingest runs at deploy time or on an upload, not per request. Nothing in the request path writes to the vector store.

data/kb/*.md ─parser─> records ─kb-db import─> database ─embeddings─> Chroma
                                    ^
                        an upload writes here too

The knowledge base lives in a database, not on the server's disk

It was one CSV that the server read off its own filesystem. That works for one machine and stops working for anything else: two servers hold two different files, a redeploy that does not preserve the volume silently reverts the corpus, and an upload made on one instance is invisible to a request served by another. None of those raise — they answer, from the wrong records.

Records live in kb_records, addressed by DECIPLE_DATABASE_URL. SQLite is the default so a fresh checkout works with nothing provisioned; it is still one file on one machine, so more than one server means pointing that URL at Postgres, which is a deployment change and not a code change. That is the only reason the records go through SQLAlchemy rather than raw SQL.

The CSV is demoted, not deleted. It stays the maintained interchange format — what a mentor edits in Excel, what kb-db import reads, what kb-db export writes. What changed is which of the two the bot believes.

Rows are the same shape as CSV rows, through the same to_rows/from_rows that the CSV uses. A second encoding would start identical and drift, and the drift would be silent: both would still load, they would just disagree about what a record means.

Writes are replace, not upsert, in one transaction. Replace because that is what writing a whole file always did — a record deleted from the corpus is deleted — and an upsert would quietly retain rows a maintainer had removed, which is the one edit that would appear to work and not. One transaction because the alternative is a delete that succeeded and an insert that did not, leaving an empty knowledge base produced by a bad upload rather than by anyone deciding to.

The knowledge base is updated by upload, not by a developer

Before this, a new lesson reached members by someone dropping a file on the server and running kb-csv, kb-parse and kb-ingest in order. That is not a code change, but it is a person with SSH access — which makes that person the bottleneck for every correction the mentor wants to make. POST /kb/preview and POST /kb/commit, behind the screen at /admin, remove them from the loop.

Preview and commit are two requests. The preview writes nothing at all, so it is safe to run against anything, and it reports what the file would add, update and leave alone. An upload is the one operation here that writes to the source of truth, and a parse that quietly produced nothing — or forty duplicates — should be visible before it lands rather than after.

A commit adds and updates; it never removes. Ingest synchronises, so a record missing from the input is deleted from the vector store on the next run. A commit that wrote only the uploaded rows would therefore delete every lesson the bot already had, and the failure would surface as members being told their question is out of scope. The merge is checked for that before anything is written.

A commit does three things or it lies. It writes the records, re-embeds, and rebuilds what the running server serves. The database is the source of truth but nothing reads it per request; the vector store is what retrieval queries; and the retriever holds the whole collection in memory from construction. Stop after the write and a mentor is told the upload worked while the bot goes on insisting it has never heard of the lesson.

Then it rewrites the CSV, which is the fourth thing and the only one a member never sees. kb-db import rebuilds the database from data/kb/, so a file left behind by an upload makes the next import an offer to delete the uploaded record — the mentor's work disappearing at a moment unconnected to the upload that added it. The whole merged set is written, exactly as kb-db export would, so the file and the database say the same thing.

It is written last, after the records are stored and the rebuilt retriever is already serving them, and a failure to write it does not fail the request: the upload has reached members by then, and reporting it as failed would be the one thing less true than saying nothing. The response carries the problem as an issue instead, naming kb-db export as the fix.

Uploads may be .md, .txt or .csv. The shape is sniffed from the content rather than the extension, so the parser's five record shapes all work whatever the file is called.

Neither endpoint takes a token. Anyone who can reach the server can change what members are sent. That was a deliberate choice for the demo; a public deployment needs something in front of these routes.

Design decisions

The model never produces KB content

The knowledge base already contains a human-written summary and an exact link for every record. Both are copied verbatim by application code.

The model's only outputs are a chosen record ID, a short empathetic reply, and a needs_handover boolean.

The reply runs to at most three sentences — validate.SENTENCE_LIMIT, enforced by trimming rather than by asking the prompt nicely. It was one sentence until 2026-09-04, which turned out to be enough room to greet somebody and none to say why the lesson suits them; the second and third sentences do that job from the "answers" and "helps when" lines, which are the only things about a lesson the model is ever shown. Widening it does not widen what it may draw on.

It cannot emit a URL because it is never asked for one. This makes invented links structurally impossible rather than something a downstream check has to catch — and a check that has to catch them is a check that will eventually miss one.

Implemented more strictly than that. The prompt does not contain the link or the summary at all. format_candidates renders each candidate as its record id, lesson title, the question the lesson answers, and its "when this helps" entries — and stops there. The model is not instructed to avoid URLs; it is never shown one, so there is no material to invent from.

Leaving the summary out costs something real. It is the fullest description of what a lesson covers, so withholding it makes the choice measurably harder than it needs to be. What it buys is that the model has never read the prose it would otherwise be tempted to paraphrase into its sentence, and a paraphrased summary presented as the lesson's own words is the failure this whole design exists to prevent.

Chunks are split inside a record, never across one

Each KB entry is self-contained: question, lesson title, link, summary, use cases. A generic text splitter would eventually cut between a summary and its link, pairing the right explanation with the wrong URL — the worst failure this system can produce, a confident and plausible answer pointing somewhere else. So chunking never crosses a record boundary, and the link is carried in metadata rather than in the chunked text, which makes that mispairing impossible by construction rather than unlikely.

Inside a record it does split, since 2026-08-18. Each record produces one chunk per retrievable facet — its question plus title, then each "when this helps" entry with the title prepended — stored as {record_id}#{n}. 76 records become 274 chunks with a median of 91 characters, down from 281.

The reason is score compression. Embedding a whole record averages its question, its title and all its use cases into one vector, and the average is blurrier than any of its parts: whole-record chunks put the top hit a mean of 0.038 above the runner-up, with the closest pair separated by 0.001. A member's question resembles one situation, not the blur of four, so retrieval scores each chunk separately and takes the record's best-matching chunk as its score — max, not mean, since averaging back would undo the split.

Measured over 256 gold questions, this lifted cosine recall@1 from 0.605 to 0.668 and widened the mean top-1 gap from 0.038 to 0.051.

Each use case is then stored twice — once with the lesson title prepended, once bare. That looks redundant and is not. The title anchors a generic situation ("You are travelling for work") to a topic, and dropping it costs recall@1 0.770 → 0.668. But the title is also vocabulary, and on a query that is pure situation with no topic words it competes with the words that matter. Keeping both copies lets each query match whichever framing suits it, and max pooling means the weaker copy is ignored rather than averaged in. Storing both beat titled-only on every metric: recall@1 0.742 → 0.770, recall@3 0.863 → 0.891, MRR 0.819 → 0.840.

This also settles a question worth recording: a granularity sweep either side of this point — use cases merged in pairs, all use cases in one chunk, long cases split at clauses, question and title as separate chunks — found nothing better. Coarser is clearly worse (all-in-one drops recall@1 to 0.633).

Every use-case chunk carries its lesson title. Bare use cases are generic situation text — "You are travelling for work" — that matches half the corpus; the title is what anchors a chunk to a topic. Dropping it costs recall@1 0.668 → 0.605, i.e. the entire benefit of chunking.

Retrieval returns records, never chunks: the reply layer still answers with one lesson and one link, so which chunk won stops at the retriever.

Correcting the premise: there were never 300-word chunks

The 2026-08-18 check-in diagnosed the low similarity scores as a 300-word chunk size packing several questions into one chunk. Two halves of that, checked against the code:

  • Chunk size. Chunks were never 300 words. Before this change a chunk was one whole record, median 281 characters — roughly 45 words. The 300 is the character median read as a word count.
  • Several questions per chunk. Never happened. Records split on Q:, which appears exactly once per lesson, so a chunk held exactly one question. The ambiguity was real but it came from between chunks, not within them: 76 records averaging one vector each, over lessons that genuinely overlap.

The conclusion survived the premise being wrong. Chunks were too coarse and making them smaller helped — recall@1 0.645 to 0.742 — but for a different reason than "too many questions in one chunk". Averaging a record's question, title, and four use cases into one vector produces a blur that matches everything weakly. The fix is one vector per idea, not fewer questions per chunk, and that distinction is why the fix is a per-facet split rather than simply a smaller window.

Why not a recursive character splitter

RecursiveCharacterTextSplitter is the default answer to chunking in most RAG material, so it was measured rather than dismissed. It splits on a hierarchy of separators — paragraph, then line, then space, then character — taking the largest that fits and recursing into anything still over chunk_size, with chunk_overlap characters repeated across the seam to avoid cutting an idea in half.

All variants below are dense-side only, max-pooled to records, over the same 256 questions (chunking-strategy experiment in MLflow, 2026-08-18):

Variant chunks median chars recall@1 recall@3 MRR margin
Whole record (pre-08-18) 76 281 0.645 0.809 0.742 0.043
Per-facet (ships) 274 91 0.742 0.863 0.819 0.054
Recursive 500/50 76 281 0.645 0.809 0.742 0.043
Recursive 256/32 125 186 0.656 0.820 0.750 0.044
Recursive 128/16 230 86 0.715 0.875 0.806 0.056
Recursive 64/8 476 49 0.703 0.852 0.786 0.052
Recursive 128/16, summaries included 415 94 0.723 0.871 0.813 0.055

Three things this settles.

At standard settings it is a no-op. Recursive 500/50 returns 76 chunks and reproduces the whole-record baseline to three decimals, because every record is already shorter than the chunk size — there is nothing for the splitter to cut. Adopting it without shrinking chunk_size would have changed nothing at all.

Size is the lever, not the algorithm. Recall climbs as chunk_size falls, under either strategy, until 64 characters where it turns over. That turnover is the floor: below roughly one sentence, a chunk stops being a retrievable idea.

Per-facet wins, but not on every metric. It leads on recall@1 by 2.7 points and MRR by 1.3, because its cuts land on semantic boundaries the source already provides and every chunk keeps its lesson title. Recursive 128/16 leads on recall@3 by 1.2 points — three questions out of 256. The split decision is worth stating plainly rather than rounding away: per-facet is better at ranking the right lesson first, recursive-128 is marginally better at getting it into the top three at all.

Recursive chunking was measured and not adopted. It loses recall@1 by 2.7 points and MRR by 1.3, and wins recall@3 by 1.2 — three questions out of 256. Losing two metrics of three to win a third by three questions is not an improvement, so the splitter and its configuration were removed rather than left in as a switch nobody would turn.

Per-facet stays because its cuts land on boundaries the source already provides, so a chunk is always a whole thought and always carries its lesson title, where a 128-character window cuts wherever it happens to land.

This was the one open item from the 2026-08-18 check-in, where recursive chunking was picked before either option had been measured. The numbers above are the answer; if the reply layer later turns out to ignore rank order inside the top 3, recall@3 becomes the only metric that matters and this is worth re-running.

What recall@1 does not measure

Exact-link scoring has a blind spot this KB walks into. Four lessons — Decision Fatigue, Understanding Our Triggers, Why Are You Eating?, Protecting Your Lifestyle — all cover negotiating with yourself over food. A member asking about that is well served by any of them; recall@1 calls three of the four a miss.

So 0.74 is a floor on quality, not a measurement of it, and the true figure cannot be recovered from link equality alone. Closing that gap needs a grader that can read a lesson and decide whether it answers the question — the "LLM as a judge" approach agreed in the 2026-08-18 check-in. That is Week 2 work. Week 1 makes no LLM calls and needs no API key.

Two things are worth measuring when it is built: how often a lesson the gold set calls a miss is actually a good answer, and how often the grader rejects a lesson the gold set calls correct. The second number is what says whether the first can be trusted.

The gold set has two tiers, and they disagree

The first 256 gold questions are one-liners. Real members do not write one-liners - they write three sentences, half of which is context the retriever does not need, with the question buried at the end. So 126 long-form questions were added (median 133 characters, up to 229) and the two tiers are scored separately, because averaging them hides the result:

tier n cosine r@1 cosine r@3 reranked r@1 reranked r@3
short one-liners 256 0.766 0.887 0.820 0.941
realistic messages 126 0.603 0.778 0.595 0.786

The headline 0.82 is measured on a distribution nobody sends. On messages shaped like real ones it is 0.60. That gap is the most decision-relevant number in the project: it is the difference between a system that looks ready and one that is.

Worse, reranking - worth +5.4 points on short queries - is worth nothing on long ones, scoring level with plain cosine. The obvious explanation was truncation, and it is wrong: zero of 1260 long-question pairs exceed the cross-encoder's 256-token limit, and raising the limit changes the numbers not at all. The real reason is distribution. ms-marco-MiniLM-L-6-v2 is trained on short web-search queries; a rambling first-person message is not what it has ever seen.

Two consequences for Week 2. A reranker trained on conversational input, or one large enough to generalise, is the obvious next experiment. And the reply layer cannot assume the top hit is usually right - at 0.60 it is right three times in five, which is an argument for surfacing three lessons and for having a genuine "I am not sure" path.

The caveat that outranks all of this: both tiers were written by the same author as the code. They are aimed at real lessons and validated against real links, but they are not real member messages, and no synthetic set can settle what real traffic will look like.

Reranking: the shortlist is re-read, not just re-scored

Every method above compares vectors built separately for the question and the passage. That is what makes searching 461 chunks affordable, and it is also a hard ceiling: the two texts are never seen together, so nothing can notice that "my head has been pounding" and "...or have a headache during the first week" are the same complaint. They embed near each other, but not near enough.

A cross-encoder reads the pair as a single input and scores it directly. It is far more accurate and far too slow to run over the whole store — so it runs over a shortlist of 10.

Ten is not arbitrary. Measured on 2026-08-18: cosine recall@10 was 0.980, against recall@3 of 0.887. The correct lesson was already being retrieved for 98% of questions and merely ordered badly, which made reranking the obvious next move and put a ceiling of 0.98 on what it could deliver.

Re-measured 2026-09-06, and the ceiling is lower than that. Over all 382 gold questions, embedding the query the request path actually sends, cosine recall@10 is 0.921 — not 0.980. The older figure predates both the larger gold set and search_text, and it is the number the pool size was argued from, so it is worth restating rather than quietly carrying forward.

Ten survives the re-measurement anyway, for a different reason. Widening the pool was measured directly, as reranked recall rather than cosine reach:

RERANK_POOL time/query rerank recall@3 recall@4
10 87 ms 0.856 0.880
20 143 ms 0.861 0.880
30 216 ms 0.856 0.869

Twenty buys half a point at rank 3 for 55ms, and thirty is worse than ten. A deeper pool hands the cross-encoder more distractors, and they displace right answers about as often as the extra reach rescues them. So the lever is not how many candidates the reranker reads — it is how many the model is shown, which is TOP_K, and that moved from 3 to 4 on the same measurement: 0.856 to 0.880, the cheapest 2.4 points available, costing four lines in a prompt.

recall@1 recall@3 MRR
Cosine alone 0.766 0.887 0.818
Cosine + cross-encoder over top 10 0.820 0.941 0.876

31 questions were promoted to rank 1; 17 were demoted from it, for a net gain of 14. The demotions are the honest cost of the trade and the first place to look if a larger reranker is ever tried — ms-marco-MiniLM-L-6-v2 is 88MB and trained on web search, not on health coaching.

Cost is 66ms per query on CPU, and it scales with the shortlist rather than the corpus: growing the KB tenfold does not make reranking slower.

The score it returns is a logit, not a similarity. It is unbounded, often negative, and comparable only within a single query's shortlist. It must never be thresholded against the cosine scale.

What a cosine score of 0.74 actually means

Cosine on bge-small-en-v1.5 does not span 0 to 1. Measured directly:

pair cosine
identical text 1.000
genuine paraphrase (a good match here) 0.673
same topic, different lesson 0.497
unrelated English sentence 0.433
gibberish 0.462

The usable band is roughly 0.43 to 1.00. Two unrelated English sentences already score 0.43 because the model packs all English into a narrow cone, and 0.9+ requires text that is nearly word-for-word identical — which a member's question never is.

So a top hit at 0.74 is not a weak match, and chasing 0.9 would mean optimising for members who paste lesson text back at us. The number that carries information is the margin to the runner-up, which is why it is logged per run. Cosine's margin is 0.060; the cross-encoder's is 4.959 on its own scale. That separation, not the absolute score, is what a confidence threshold can eventually be built on.

BM25 scores are normalised to 0-1

Raw BM25 is unbounded — 4.9 to 29.5 on this corpus — which makes it unreadable next to a cosine 0.71 and makes a single confidence threshold across the two impossible to state. Scores pass through s / (s + 11.0) on the way out of the retriever, where 11.0 is the measured median top-1 raw score, so 0.5 means "a typical best match".

The transform is monotonic, so every recall and MRR figure is unchanged by it — verified by re-running the eval before and after. Min-max scaling over the result set is the obvious alternative and is worse: it pins the top hit to exactly 1.0 on every query, including the ones where nothing matched, which destroys the signal the normalisation exists to provide.

Dense and sparse search different text — deliberately

The question, lesson title, and "when this helps" entries are embedded. Summaries are not: every summary in this KB shares generic vocabulary — "low-carb", "blood sugar", "cravings" — so embedding them pulls unrelated records toward every query. The use-case entries, by contrast, are written as descriptions of member situations, which is what a member's question actually resembles; titles are short and distinctive.

BM25 does index the summaries, because the same argument runs the other way for sparse retrieval: idf discounts corpus-common terms automatically, so a summary contributes only its rare terms ("keto flu", "wine") — terms members actually type and nothing else in the record contains. The measurements behind this split are in the eval section below.

Link and summary are stored in metadata either way: what the reply layer serves is never dependent on what retrieval searched.

What gets embedded is the question, not the whole post

A community post is mostly not a question. It opens with a greeting, gives six sentences of context, and ends with the one line actually being asked. Embedded whole, that averages out: the same coffee-and-fasting question that retrieves "Fasting Lever: Basics" on its own retrieves "The Importance of Self Care" behind a paragraph of preamble, and does it at a cosine of 0.75 — comfortably above the floor, so nothing flags it. The model is then asked to choose between three lessons about the mood of the post rather than its question.

So when the message contains a question, that is what gets embedded. A message that is only a question — most of them — is unchanged. A message with no question mark falls back to the whole text, because the alternative is guessing which sentence carries the intent. Retrieval is the only thing this narrows: the model, the reply and the handover email all still see the member's words in full.

The failure it introduced, found 2026-09-06. A question sentence can be anaphoric — the topic sits in the context, and the question refers back to it with a pronoun:

My husband does the fasting thing and I do the low carb thing and we keep arguing about which is better. Is there something that explains how they actually interact?

Extracting the question is correct and throws away "fasting" and "low carb", leaving a sentence whose only topic word is "they". Cosine on the whole post scores 0.761; on the question alone, 0.563 — below SCOPE_FLOOR. A lesson the retriever had at rank 1 was refused as off topic, and the member was told their question was out of scope.

The retry is on the symptom, not the sentence. No lexical test catches this: the sentence is well formed, has four content words, and reads as a perfectly good question. What is checkable is the outcome — so when the confidence lands below the floor and the extracted question differs from the message, the whole message is scored too, and the better of the two is used. That costs one extra dense query, only on the branch that was about to refuse somebody, and it cannot rescue a genuinely off-topic post because a post about parking scores badly either way.

Safety runs before retrieval

Not for ordering tidiness. Because retrieval scores high on exactly the questions that most need a human.

The KB covers heart palpitations and kidney stones. A member writing "my heart has been racing since I started keto" matches the Electrolytes lesson almost exactly. The retrieval score therefore pushes toward auto-replying on precisely the messages that should not be auto-replied to.

A confidence threshold cannot catch this — confidence is high, and it is high for a legitimate reason. The only place to catch it is before retrieval runs.

The gate is a word list, not a model

Three reasons, in order of weight. It cannot fail on a timeout, because it makes no network call. Every decision traces to one listed phrase, so a disagreement about a classification is settled by reading a line rather than by re-running anything. And it is testable case by case — which matters more here than anywhere else in the system, because the branch under test is the one that handles someone saying they want to die.

Phrases are long on purpose: "kill myself", never "kill". This knowledge base is about food, and the short forms are ordinary speech in it — "this diet is killing me", "I'm dying for chocolate". A gate that fires on those is a gate the team learns to ignore.

Crisis and sensitive are not the same trade. On the sensitive branch a false positive costs one unnecessary email, so it is written broadly. On the crisis branch a false positive costs a member their actual answer and hands them a hotline they did not need, so it is written narrowly and one exception is applied: a crisis phrase preceded by avoidance language does not count. This was not theoretical. A real gold question — "is there proper guidance on doing both without hurting myself?" — classified as a crisis, which would have answered a question about fasting with a suicide hotline. Negation flips it back, because "I can't get through a day without hurting myself" is the same words meaning the opposite thing.

The medication rule asks what is being asked, not which drug is named. Naming one used to be enough, and the cost was measured: three lessons became unreachable. "Weight Loss Medications", "What You Need to Know About Ozempic and Mounjaro" and "Dr. Tro on GLP-1 Medications" each exist to answer a question the gate stopped before retrieval ran, so the member the lesson was written for was the one member who could not be given it. The same list screens the model's reply, so a reply naming the drug was blocked on the way out too — fixing only the input would have served the same question one time and handed it over the next, depending on whether the model happened to name the drug.

So the rule holds treatment framing — "my dose", "stop taking", "come off my", "prescribed" — and not drug names. A bare drug name reaches retrieval, where a lesson answers it or the model reports that none fits and it becomes a handover one step later; the email still arrives, after the knowledge base has been given its chance. Questions about a member's own prescription are unchanged.

Two verbs were tried and withdrawn against the gold set, which is the reason they are worth recording. Bare "come off" read "I want to come off Mounjaro" as a treatment decision, but the mentor wrote a lesson answering exactly that, so the possessive in "come off my" is what carries the meaning. And "should i take" stopped "what precautions should I take before fasting" — not a medication question at all.

Measured over the 382 gold questions, 2 (0.5%) then stopped before retrieval, both about bingeing, down from 8 (2.1%). Unreachable lessons went from four to one — "Tales From the Scale", stopped by the eating-disorder rule. Re-measured after that rule was narrowed on 2026-09-18: 0 of 382 stop before retrieval, and no lesson is unreachable. No crisis false positives remain.

The eating-disorder rule was narrowed to compensation

"binge", "binged", "bingeing", "binging", "hiding food" and "hide food" left the rule on 2026-09-18. It is the medication rule's lesson applied a second time: what decides is whether the knowledge base can answer, not whether the subject sounds alarming.

The mentor wrote the answers. "Our Food Relationship Assessment" is a self-check built around sneaking food, hiding wrappers and shame, and "How To Recover When You've Been Off Plan" is about stopping the binge-restrict cycle. Retrieval matches them at 0.70 to 0.81. A member writing the sentence those lessons were written for was getting an email and a wait.

The trade is real and runs the other way from the rest of this document, which prefers a needless handover to a wrong answer. Here the "needless handover" was being sent to the member least likely to ask twice — someone who had just admitted something they are ashamed of. The reply they got named medication and diagnoses, neither of which they had mentioned.

What still stops. Physical compensation — purging, vomiting, laxatives — and prolonged restriction: "starving myself", "not eaten in days". No lesson can answer those, and they carry acute risk of their own. Bingeing described with compensation still escalates, through the model gate rather than the word list: "I eat a huge amount at night and then skip food for two days" reads as sensitive on the compensation. So does eating to the point of being physically sick, most nights — measured, and the one case in that family with no compensation to catch it.

The known gap was paraphrase. A phrasing nobody listed gets through: the confidence threshold and the model's own needs_handover flag stood behind the rules for the sensitive cases, but neither helped on the crisis branch, because nothing later in the pipeline was looking for one. "Everyone would be better off if I just wasn't around anymore" reached retrieval like any other question.

The second gate: a model, behind the rules

intent.IntentGate closes it, and where it sits is the whole design. It runs only when the rules found nothing, which buys three things at once. A crisis the rules recognise is still answered instantly and offline, so the branch that matters most cannot be delayed or lost by a model that is slow, rate limited or down. Every verdict the rules reach still traces to one listed phrase. And the model is asked only about messages that already look ordinary — where its judgement adds something, and where being wrong costs least.

It fails open. A timeout or an error returns ok and the question carries on to retrieval, which is exactly what happened before the gate existed: the gap it leaves on failure is the gap that was already there. Failing closed would turn a Groq incident into every member being told to phone the office.

The trade is one extra LLM call on questions that turn out to be fine, which is most of them, and it is recorded in its own intent trace span so a handover can be traced to whichever gate produced it. The two are deliberately distinguishable: model_intent is not a rule name, because a listed phrase and a model's reading of a sentence warrant different amounts of trust.

Verified end to end against the running API: "everyone would be better off if I just wasn't around anymore" and "should I double up on the tablets my doctor gave me" both reach retrieval under the rules alone, and become crisis and handover respectively with the gate in front of it. "How much salt do I need on keto" is unchanged — lesson, 0.78.

The gate reads what members write; a second screen reads what we write back

validate.check screens the model's sentence before a member sees it. That is a different job from the gate, and for a while it was doing it with the gate's rules — the inbound classifier, run over the bot's own words.

Measured 2026-09-06: 3 of 60 traces lost a correct answer to it. Every one had already chosen the right lesson. Each became a handover a person then had to answer. All three tripped on the same word:

This lesson is designed for moments when you catch yourself justifying a binge with "I'll start over Monday".

The eating_disorder rule exists to catch a member disclosing their own situation — "I binged last night". Read as vocabulary it is simply a subject the knowledge base is written about, and a lesson about bingeing cannot be introduced without the word. The member's message had already passed the gate; the topic came from the lesson, not from them.

So the two screens now differ, and the split is by shape of phrase, not by severity. A rule stays on the outbound screen when a sentence containing it is a sentence making a medical claim, whoever wrote it: crisis phrasing, acute symptoms, medication instructions. "Chest pain is normal in the first week" and "you can stop taking that" are advice however they are framed, and still fail. The three rules that only name a subject — eating disorder, diagnosed condition, pregnancy — no longer do.

Inbound classification was untouched by that change. It changed later, on 2026-09-18 and for the same reason read from the other side: a member who writes "I binged last night" now reaches the lesson written for them. See "The eating-disorder rule was narrowed to compensation" below.

This is the one place in the design where a guardrail was deliberately loosened, so it is worth being explicit about the trade. The governing trade-off below prefers a needless handover to a wrong answer — and this change still honours it, because what was removed was never protecting against a wrong answer. It was rejecting correct ones for containing the name of their own subject.

The empathy sentence is discarded on the handover branches

Discovered by running it. Asked "what time does the front desk close?", the model correctly returned no lesson and set needs_handover — and then wrote "let me help you find the front desk hours." Nobody is going to.

The sentence is written to introduce a lesson. With no lesson attached it becomes a promise the system cannot keep, so the handover and scope replies use fixed text and drop it. The model's judgement is used; its prose is not.

Every reply carries the disclaimer

A compliance requirement, and the one rule with no exceptions: every branch returns it, including the ones that never reach the model. A crisis reply, an off-topic refusal and a served lesson all carry the same sentence.

It is a separate field, not glued onto the message. AskResponse returns message and disclaimer apart, and the surface renders them apart. Blended into the prose it becomes something a future prompt change can swallow, and something the model appears to have said; kept separate it is applied by the code that assembles the response, on a path the model has no say in. Three tests hold this down — that every branch carries it, that the two fields never merge, and that reconfiguring the wording changes every branch at once.

It has a working default, unlike DECIPLE_CRISIS_RESOURCES, which is left blank on purpose because the right hotline depends on where a member is and a wrong number is worse than none. "Call your local emergency number" is true everywhere, so a deployment that never configures it is still compliant rather than silently unprotected. The wording belongs to whoever owns compliance; setting it empty removes it from every reply, which is why a test asserts that cannot happen by accident.

Four exit points, each with a defined destination

Low confidence, off-topic, model-flagged uncertainty, and failed verification are four different conditions, discovered at four different stages. Each gets its own exit rather than collapsing into one shared failure path, because the right response differs.

Off-topic questions get a canned scope reply rather than a handover email — nobody needs to follow up on a request for a gym recommendation.

On-topic questions the KB does not cover reach the member honestly and email nobody. That is a deliberate reversal, and it cost something. The original reasoning was that a missing answer is a gap worth knowing about, which is true. What made it untenable was the volume: a probe of sixteen edge-case messages produced eleven emails, and nine came from this one branch — a greeting, "thanks!", "test", a single "why", a keysmash and a spam link among them. None of those is a gap in the knowledge base. An inbox that is mostly noise stops being read, and what it stops being read for is the sensitive handover sitting in it, which is the one message here that cannot wait.

The gap signal is not lost, it moved. needs_email is stamped on every root span as False rather than omitted, alongside outcome and reason, so "show me every question the courses could not answer this week" is still one MLflow query — it is just no longer sixty emails. Two guards keep the volume down before that branch is ever reached: the scope floor turns away off-topic questions, and the substance check turns away messages that are not questions.

What still emails: a message the safety gate judged sensitive, and a reply that failed validation. Both mean a person is needed.

The handover email goes over Resend's API, not SMTP

The email is the same email. What changes is that every send comes back with an id, and the dashboard then says whether it was delivered, bounced or marked as spam. An SMTP handshake that succeeds tells you a relay accepted the message, not that anybody received it — and the whole point of a handover is that a person actually picks it up.

Composing and sending are separate functions. Composing is pure: the same answer in, the same bytes out, so what the team will read is asserted in a test with no mail server anywhere near it. Sending is the part that talks to the world and is allowed to fail.

An unconfigured mailer logs instead of raising. A missing API key is a deployment that has not finished, not a bug, and the member has already been told a human will follow up. Raising there would take a reply that worked and break it over a side effect. The composed email goes to the log at WARNING, so nothing is lost and the gap is visible.

Nothing here is called from the request path. The request already spends 1.9 seconds in the model, and a second network round trip in front of a member's reply buys them nothing, so the API layer sends this after the response has gone out.

The email carries the member's question verbatim, the reason for the handover, the confidence score with the floor it was judged against, and the ranked candidates with their scores and links. On the branch where a clinical question stops at the safety gate there is no score and no shortlist, because retrieval never ran — the email says so in those words rather than printing a zero that reads like a real measurement.

Not yet carried: the community post id. The team asked for it so a handover can be traced back to the post that caused it. Nothing supplies one today — /ask is called by the member chat, which has no posts. It arrives with the Disciple adapter, which is the surface that has them.

One threshold, not two

The design above called for two confidence checks: one to spot off-topic questions, and one to spot on-topic questions the knowledge base cannot answer. Only the first exists, because only the first is measurable.

The second was measured and abandoned. Over 374 gold questions, cosine scored a median 0.729 where the right lesson was retrieved and 0.682 where it was not. The distributions almost coincide: a floor catching half the misses throws away a fifth of the good answers. There is no number that separates "the KB has this" from "the KB does not", so no threshold can do that job. The model's needs_handover flag does it instead — it reads the candidates rather than scoring them, which is the one thing a scalar cannot do.

The first works, on cosine and not on the rerank score. The cross-encoder orders candidates far better, but its output is an uncalibrated logit and its ranges overlap almost completely: on-topic ran −11.1 to 7.9, twenty deliberately off-topic probes ran −11.3 to −9.0. "There's a wedding coming up, can I drink?" — a real gold question with a real lesson — scored −11.141, below eighteen of the twenty probes. Cosine is bounded and behaves: on-topic 0.511 to 0.944, off-topic 0.413 to 0.608. Hence SCOPE_FLOOR = 0.58, which catches 80% of the probes and costs 3 of 374 real questions.

So the two scores are used for different jobs on purpose — the reranker orders, cosine judges. That costs one extra dense query per request, which is noise beside a 1.9 s model call.

The floor is calibrated against twenty hand-written probes, which is thin evidence. It is deliberately conservative for that reason, and because it is not the last line of defence: an off-topic question that clears it still reaches the model, which reliably declines. What the floor actually decides is whether a rejected question generates an email or a canned reply.

The governing trade-off

Two failure modes with unequal costs:

  • Saying "I don't know" when it could have helped — mildly annoying.
  • Confidently giving a wrong or inappropriate answer — much worse.

Every branch above trades in favour of the first.

Retrieval approach

Cosine similarity alone is weak on rare and specific terms. This KB uses distinctive phrasing that could disambiguate its semantically overlapping lessons, and that signal is exactly what a dense embedding tends to smooth away.

So BM25 will be measured alongside cosine, and the two fused with Reciprocal Rank Fusion — rank-based voting, which avoids having to normalise two different score scales against each other.

Baseline cosine numbers come first, so the fusion's effect is measurable rather than assumed.

Method recall@1 recall@3 MRR top score margin
Cosine 0.71 0.85 0.77 0.727 0.051
BM25 0.46 0.63 0.54 0.539 0.058
RRF fusion 0.65 0.81 0.72 0.032 0.002
Cosine + rerank 0.74 0.89 0.81 −2.909 4.234

top score is the mean score of the rank-1 hit and margin the mean gap to rank 2 — both logged per run, so a chunking change can be judged on separation and not only on recall. BM25 is on a normalised 0-1 scale (see below); the RRF column is an ordering, not a confidence, and its 0.032 should not be read as one.

Measured 2026-08-18 with BAAI/bge-small-en-v1.5, per-facet chunks and the query instruction prefix, against the 256 short-form questions of the 382-question gold set (data/eval/questions.jsonl), covering all 64 linked lessons.

The other 126 are long-form messages and are not in this table. They score far lower, and they are scored separately for the reason given in "The gold set has two tiers, and they disagree" above: averaging the two hides the only number that decides whether this is ready. Running kb-eval over the whole file produces that average — 0.71 cosine recall@1, 0.74 reranked — and it should not be quoted as a headline, because it describes a question distribution that is half artificial and half real.

This table now scores all 382 questions together, where it previously ran against the 256 short ones alone and read 0.82 / 0.94 / 0.87 for reranking. Nothing regressed: re-scored on its own, the short tier still returns 0.941 recall@3 exactly. The whole of the drop is the 126 long-form questions, which are a different and much harder problem — see "The gold set has two tiers, and they disagree" above, which measures the split and rules out the obvious explanation for it.

The combined figure is the honest headline, and it is the one to quote. Neither tier alone describes what a member sends: quoting 0.94 overstates the system, and quoting 0.786 ignores that plenty of real questions genuinely are short.

An earlier revision of this table is still worth recording, for the same reason. It read 0.80 / 1.00 / 0.89 for cosine — measured against 15 questions, where a single question moved recall@3 by 0.07 and a perfect score meant "no counter-example in fifteen tries". At 256 questions one question is worth 0.004, and recall@3 of 1.00 turns out to have been an artifact of the sample rather than a property of the system.

Getting the gold set to that size was the precondition for every other decision below: at n=15 the chunking change and the query prefix were both inside the noise, and neither could have been accepted or rejected honestly.

A stronger embedding model reversed which method wins

The table above is the second measurement. The first used all-MiniLM-L6-v2, and under it the fusion led:

Method (all-MiniLM-L6-v2) recall@1 recall@3 MRR
Cosine 0.73 0.80 0.77
BM25 0.47 0.80 0.60
RRF fusion 0.73 0.93 0.83

One question missed under all three methods: "Is it okay to have a glass of wine now and then on this diet?" The lesson that answers it is written entirely in alcohol vocabulary — "wine" appears once, as "wines", in the summary. A control query phrased with "alcohol" returned that record at rank 1 in all three methods, which located the failure in the embedding rather than in the composition or the fusion.

bge-small-en-v1.5 — same size class, stronger on paraphrase — closes that gap and takes cosine to 1.00 recall@3. It also inverts the hybrid result. RRF rewards agreement, and BM25 cannot agree about a record whose distinguishing term it never matches: the wine record is cosine's rank 1 and absent from BM25's list entirely, so records both methods rank mediocrely outvote it out of the top 3.

The hybrid design was left on probation at that point, because fifteen questions cannot separate 0.93 from 1.00 with any confidence.

Resolved 2026-08-18 on 256 questions. Fusion has the best recall@3 (0.88 vs cosine's 0.86) and ties cosine on MRR, but loses recall@1 (0.71 vs 0.74). The probation ends in a split decision rather than a winner: fusion is better at getting the right lesson into the top three, cosine is better at putting it first. Since the reply layer chooses among a handful of candidates, recall@3 is the metric that matches how retrieval is actually consumed, and fusion keeps its row on that basis — not on the assumption it started with. (TOP_K is four since 2026-09-06, so recall@4 is now the closer match; it did not change which method wins.)

How the composition was chosen

Measured with all-MiniLM-L6-v2, before the model change above. Variant C's composition is retained under the new model: the asymmetry is a property of this corpus, not of the embedding.

Measuring before assuming paid off, in two stages. On the first ingest (dense and sparse both searching question + use cases), fusion lost to the cosine baseline: BM25 was weak on paraphrases — members don't reuse the KB's vocabulary — and its votes dragged correct cosine top-1s down (recall@1 0.67 → 0.53). Fusion had failed to earn its row.

The miss analysis showed why: the exact terms members do use ("keto flu", "wine") exist in the KB, but only in the summaries — the one field neither method searched. Three compositions were then measured (each an MLflow run named experiment-* in the retrieval-eval experiment):

Variant (fused) recall@1 recall@3 MRR
A — dense = sparse = question + use cases 0.53 0.87 0.68
B — A + lesson title in both 0.73 0.87 0.80
C — B + summaries in BM25 only 0.73 0.93 0.83

Variant C is what ships. Titles help both methods (short, distinctive, none of the shared vocabulary). Summaries are added asymmetrically: embedding them would pull every record toward every query, but BM25's idf discounts corpus-common terms automatically, so a summary contributes only its rare terms — exactly the signal BM25 exists to capture. With that split, fusion beat the cosine baseline on every metric — by measurement, not assumption. The stronger embedding model then changed that answer again, which is the section above.

Expected retrieval characteristics

The KB deliberately covers overlapping ground. Several records address internal negotiation over food:

  • Decision Fatigue
  • Understanding Our Triggers
  • Why Are You Eating?
  • Protecting Your Lifestyle

A question about that will score all four nearly identically, because they are genuinely about the same thing from different angles.

The prediction, then: recall@3 should be strong while recall@1 is mediocre.

That gap is the empirical argument for letting the model choose among three retrieved records rather than always injecting the top-ranked link. If recall@1 turns out to be strong, the model's role in selection shrinks accordingly — so this number is worth measuring before the reply layer is designed around it.

Measured outcome (256 questions, 2026-08-18): the prediction holds, and holds more clearly than the first fifteen questions suggested — 0.74 recall@1 against 0.88 recall@3 for fusion. Roughly one question in four does not resolve at rank 1, and the overlapping records above are where those near-misses land.

Serving several records is therefore justified by measurement rather than by caution. (TOP_K was three when this was written and is four since 2026-09-06; see the reranking section for the measurement that moved it.) But the reply layer should be designed knowing the top hit is usually right: the model's job is to reject an obviously wrong rank-1, not to re-rank from scratch.

The 0.12 that misses even at rank 3 is the number that matters for Week 2. It is the floor on how often a handover — not a lesson link — is the correct reply, and it is the reason a confidence threshold has to exist at all.

Three evaluation tools, each answering a different question

They are not three versions of the same score. Each one is unable to answer the question the next one asks.

Tool Question Ground truth
kb-eval Does the retriever rank the right lesson highly? Gold links
kb-grade Did the member get what they should have? Gold links and outcomes
kb-judge Why was a reply poor? None — a model reads the traces

kb-eval stops short of the answer. Recall@3 says the right lesson was available to the model. It does not say it was given. Between the two sits the model, which reads the shortlist and picks one record or none.

kb-grade closes that gap, and its point is the blame split. A wrong answer has two causes needing opposite fixes. If the expected lesson was never in the shortlist, the retriever is at fault and no prompt will help. If it was sitting there and the model handed over anyway, the retriever is fine and the prompt is the thing to change. One accuracy number without that split says something is wrong and nothing about where.

Not every question should be answered with a lesson. A member asking what a screening costs, or whether there is personal training, is asking something the courses do not cover, and the right reply is a person rather than the nearest lesson. Those questions carry expected_outcome in the gold file — data/eval/handover.jsonl — and are scored on the outcome instead of a link, because a correct refusal has no link to compare and would otherwise be marked wrong for serving nothing. They are blamed on the model when they miss: declining is the model's own call, since it is shown a shortlist and returns NONE or does not.

It gets the rank by querying the retriever a second time rather than reading it out of the trace. Retrieval is deterministic and costs milliseconds against the model's two seconds, so the duplicate query is nearly free — and it is the only way to get a rank for questions that never reached retrieval. A question refused at the scope floor still has an answer to "would the right lesson have been there?", and that answer is the whole case for moving the floor.

kb-judge asks why, which neither of the others can. A trace holds what the model was shown and what it picked, so the reason is in there; there are just more traces than anyone will read by hand. One LLM call per trace, judging the decision against the candidates that were on the table at the time.

Its verdicts are a fixed set, not free text. A paragraph per trace explains one reply and aggregates into nothing; four labels turn a hundred traces into "sixty-two served well, nineteen over-refused". The labels separate because the fixes differ: wrong_candidate and over_refused are prompt problems — the right lesson was in front of the model and it passed it over or refused anyway — while correctly_refused is a knowledge base problem, and no prompt fixes a lesson that does not exist.

The judge never sees a safety stop. Questions that hit the crisis or clinical gate are excluded before any call is made. The judge would be reasoning about whether refusing a medical question was too cautious, and a model's opinion there is not evidence anyone should act on — it is a plausible sentence arguing to weaken a guardrail. The gate's correctness is a matter for the safety tests and a person, so those traces are counted and set aside.

Knowledge base format

The mentor's files were expected to use two layouts; the real corpus turned out to contain five record shapes, and one file mixes two of them — so shape handling is per-record, not per-file (the full inventory is documented in kb_parser.py):

  1. Q:/A: block — one field per paragraph, * bullets, markdown links
  2. Q:/A: inline — every field on the A: line, bare URLs, • or unmarked use cases
  3. ## N. Question headings (FAQ.md)
  4. A markdown table of contents — its rows don't become records; they enrich existing lesson records, joined on the lesson URL (title matching scores 0 of 56 against the Lesson: fields)
  5. ## N. URL crawler entries — no question text, so one is derived from the URL slug

Records are split on the Q: line where present, because it is the only boundary marker every Q-style shape shares; within a record, the field labels (Lesson:, Format:, …) are the only reliable structure, which is what lets one field-splitter handle both the block and inline shapes.

Two hosts appear for the same lessons. community.doctortro.com 301-redirects to community.toward.health (verified live), so every link is canonicalised to the latter at parse time — the majority host in the corpus is the stale one.