Overview
A member of Toward Health's community app asks a question. The bot replies with a short empathetic note — at most three sentences — the exact link to the course lesson that answers the question, and that lesson's summary.
It is a librarian, not a clinician. It finds the resource that already answers the question; it does not answer health questions itself.
Scope boundaries
- Links come verbatim from the knowledge base. The LLM never generates a URL.
- Summaries come verbatim from the knowledge base. The LLM does not rewrite, condense, or paraphrase them.
- No diagnosis and no treatment advice. Not hedged, not caveated — not given.
- A human is involved when the bot is unsure, or when the question is too sensitive to auto-answer.
- Every reply carries the compliance disclaimer. On every branch, including the ones that never reach the model.
Status
Week 1 was the retrieval core only, and answered one narrow question:
Given a member's question, can we find the right lesson?
Week 2 built everything that stands on it — the reply, the safety gate, the member chat, the handover email, and the tools to measure whether the answers are any good.
Week 3 added the second surface and the tooling around the knowledge base: the Disciple adapter, an intent gate in front of retrieval, a one-command ingest that a running server actually picks up, and a handover email strict enough to be worth reading.
Week 1 — retrieval
| Task | Status |
|---|---|
| Repo scaffold, MkDocs | Done |
| Parse KB into records | Done — 76 records from five source shapes (kb-parse); 10 web pages fail validation for having no summary, so 66 are embedded as 461 chunks |
| LangChain + Chroma ingest | Done — cosine collection, unchanged records skipped (kb-ingest) |
| Retrieve top k + scores | Done — cosine, BM25, RRF fusion, and a cross-encoder rerank; k went from 3 to 4 on 2026-09-06, measured |
| CLI eval on a gold question set | Done — kb-eval over 382 questions covering all 64 linked lessons |
| MLflow retrieval traces | Done — one trace per question, one span per method |
Week 2 — the reply, the surfaces, the tooling
| Task | Status |
|---|---|
| Reply layer | Done — DSPy over Groq, at most three empathetic sentences (validate.SENTENCE_LIMIT), link and summary copied verbatim |
| Safety gate | Done — word list, runs before retrieval, separates sensitive from crisis |
| Off-topic floor | Done — cosine below DECIPLE_SCOPE_FLOOR gets the canned scope reply, no email |
| Handover email | Done — Resend's HTTP API, carries the question, the confidence and the ranked candidates |
| Compliance disclaimer | Done — on every reply, on every branch, defaulted so it cannot be forgotten |
| HTTP entry point | Done — FastAPI /ask, /health, /kb/preview, /kb/commit |
| Member chat | Done — React and Vite, one surface, no memory between messages |
| Knowledge base in a database | Done — SQLAlchemy, kb-db import/export/status; the server no longer reads a file |
| Knowledge base uploads | Done — upload screen at /admin, preview then commit |
| Answer eval | Done — kb-grade scores served answers, kb-judge reads the traces with an LLM |
| Disciple adapter | Moved to week 3, where it was built — see below |
| Chat memory | Deferred — the browser keeps a transcript per tab, but no history is sent to the model and each question is answered alone |
Week 3 — the second surface, and the tools around the knowledge base
| Task | Status |
|---|---|
| Disciple adapter | Built and tested, and off until it is configured — webhook endpoint, HMAC-SHA1 signature over the raw body, fast acknowledgement, and a claim so a retried delivery is answered once. Written against guesses, then corrected against Disciple's published docs, which contradicted three of them. All four DECIPLE_DISCIPLE_* values are still empty — scripts/disciple_local_test.py runs the whole loop against a stand-in so it can be seen working without an account |
| Intent gate | Done — a short model call reads what the message is before retrieval runs, so a greeting and a question are not treated alike |
| Chat transcript in the browser | Done — session storage, one tab, cleared when the tab closes; nothing is stored server side |
| Answer eval for handovers | Done — kb-grade covers the questions that should reach a person, not only the ones that should be answered |
| Knowledge base ingest from the command line | Done — npm run ingest from the repo root, and kb-reload so a running server picks the change up without a restart. Without that last step the store held 67 records while the bot still answered from 66 |
| Safety filter tuned against measurement | Done — a 46-message probe across stress eating, discouragement, side effects, medication, symptoms, disordered eating and crisis: 45 right, the one exception erring toward a person. Fixed three real misfires: a stress-eating question a member actually sent, constipation escalated as needing clinical guidance, and the eating-disorder rule sending "I hide wrappers" to the team instead of the self-assessment lesson written for it. Gold questions stopping before retrieval went from 2 to 0 |
| Uploads keep the CSV in step | Done — a commit rewrites data/kb/knowledge_base.csv from the database as its fourth step. Before that an upload left the file stale, and the next kb-db import offered to delete the uploaded record — the mentor's work disappearing at a moment unconnected to the upload that added it |
| Uploads without a token | Done — the token was removed entirely. Anyone who can reach the server can upload, so it must not be exposed publicly as it is |
| Handover email strictness | Done — a substance check turns away greetings, keysmashes and one-word messages before retrieval, and "no lesson fits" no longer emails. A sixteen-message probe went from eleven emails to one, and that one is the clinical case |
| Knowledge base on Google Drive | Researched only. Drive content is reachable and comes back in the exact shape kb_parser already reads, so no new parser is needed. Nothing is built: the bot has no Google credentials, and it would need a service account on a shared folder before Drive could be a source |
What is left
Not code, mostly. The Disciple adapter needs a portal subdomain, an API key, the bot's numeric author id and a webhook secret — and an answer to the one question in the Disciple notes that is nobody's to decide alone: whether the bot should ever reply in public to a post from someone in crisis.
The handover email needs doctortro.com verified with Resend. Until those DNS
records exist the only working sender is Resend's sandbox address, which
delivers to the account owner and silently drops everyone else — so adding a
teammate to DECIPLE_HANDOVER_TO today would look like it worked and would not.
Verified working to the account owner's address on 2026-09-27.
And the one piece of code a public deployment needs: nothing guards
/kb/preview, /kb/commit or /kb/reload. Anyone who can reach the server can
change what members are told, and rebuilding the retriever is the most expensive
thing the process can be asked to do. On a laptop that is the deliberate choice
recorded in the architecture notes. On a public URL it is the blocker.
A known limitation, measured
The bot cannot tell "the right lesson is missing" from "the right lesson is here". Asked something the courses are near but do not cover, it offers the closest lesson it has, at ordinary confidence.
Measured on 2026-09-17 by deleting each lesson in turn and re-asking its gold
question: all 64 stayed above DECIPLE_SCOPE_FLOOR, on a substitute. The floor
is one absolute cosine threshold, and the two cases overlap almost exactly —
0.580–0.880 with the lesson present, 0.580–0.777 without it. Raising the floor
to 0.75 would catch 91% of the missing-lesson cases and reject 55% of correct
answers. The cross-encoder separates them no better. Shown 12 shortlists with
the right lesson removed, the model itself declined twice.
The floor still does the job it was built for: a genuinely distant question ("what's the best pizza topping?") scores below it and gets the scope reply. What it cannot detect is missing coverage, and that needs a different mechanism — a second model call asking whether the chosen lesson answers the question, or a reply that offers the lesson as the closest thing found rather than as the answer. Neither is built, and both need their own eval.