Skip to content

Overview

A member of Toward Health's community app asks a question. The bot replies with a short empathetic note — at most three sentences — the exact link to the course lesson that answers the question, and that lesson's summary.

It is a librarian, not a clinician. It finds the resource that already answers the question; it does not answer health questions itself.

Scope boundaries

  • Links come verbatim from the knowledge base. The LLM never generates a URL.
  • Summaries come verbatim from the knowledge base. The LLM does not rewrite, condense, or paraphrase them.
  • No diagnosis and no treatment advice. Not hedged, not caveated — not given.
  • A human is involved when the bot is unsure, or when the question is too sensitive to auto-answer.
  • Every reply carries the compliance disclaimer. On every branch, including the ones that never reach the model.

Status

Week 1 was the retrieval core only, and answered one narrow question:

Given a member's question, can we find the right lesson?

Week 2 built everything that stands on it — the reply, the safety gate, the member chat, the handover email, and the tools to measure whether the answers are any good.

Week 3 added the second surface and the tooling around the knowledge base: the Disciple adapter, an intent gate in front of retrieval, a one-command ingest that a running server actually picks up, and a handover email strict enough to be worth reading.

Week 1 — retrieval

Task Status
Repo scaffold, MkDocs Done
Parse KB into records Done — 76 records from five source shapes (kb-parse); 10 web pages fail validation for having no summary, so 66 are embedded as 461 chunks
LangChain + Chroma ingest Done — cosine collection, unchanged records skipped (kb-ingest)
Retrieve top k + scores Done — cosine, BM25, RRF fusion, and a cross-encoder rerank; k went from 3 to 4 on 2026-09-06, measured
CLI eval on a gold question set Done — kb-eval over 382 questions covering all 64 linked lessons
MLflow retrieval traces Done — one trace per question, one span per method

Week 2 — the reply, the surfaces, the tooling

Task Status
Reply layer Done — DSPy over Groq, at most three empathetic sentences (validate.SENTENCE_LIMIT), link and summary copied verbatim
Safety gate Done — word list, runs before retrieval, separates sensitive from crisis
Off-topic floor Done — cosine below DECIPLE_SCOPE_FLOOR gets the canned scope reply, no email
Handover email Done — Resend's HTTP API, carries the question, the confidence and the ranked candidates
Compliance disclaimer Done — on every reply, on every branch, defaulted so it cannot be forgotten
HTTP entry point Done — FastAPI /ask, /health, /kb/preview, /kb/commit
Member chat Done — React and Vite, one surface, no memory between messages
Knowledge base in a database Done — SQLAlchemy, kb-db import/export/status; the server no longer reads a file
Knowledge base uploads Done — upload screen at /admin, preview then commit
Answer eval Done — kb-grade scores served answers, kb-judge reads the traces with an LLM
Disciple adapter Moved to week 3, where it was built — see below
Chat memory Deferred — the browser keeps a transcript per tab, but no history is sent to the model and each question is answered alone

Week 3 — the second surface, and the tools around the knowledge base

Task Status
Disciple adapter Built and tested, and off until it is configured — webhook endpoint, HMAC-SHA1 signature over the raw body, fast acknowledgement, and a claim so a retried delivery is answered once. Written against guesses, then corrected against Disciple's published docs, which contradicted three of them. All four DECIPLE_DISCIPLE_* values are still empty — scripts/disciple_local_test.py runs the whole loop against a stand-in so it can be seen working without an account
Intent gate Done — a short model call reads what the message is before retrieval runs, so a greeting and a question are not treated alike
Chat transcript in the browser Done — session storage, one tab, cleared when the tab closes; nothing is stored server side
Answer eval for handovers Done — kb-grade covers the questions that should reach a person, not only the ones that should be answered
Knowledge base ingest from the command line Done — npm run ingest from the repo root, and kb-reload so a running server picks the change up without a restart. Without that last step the store held 67 records while the bot still answered from 66
Safety filter tuned against measurement Done — a 46-message probe across stress eating, discouragement, side effects, medication, symptoms, disordered eating and crisis: 45 right, the one exception erring toward a person. Fixed three real misfires: a stress-eating question a member actually sent, constipation escalated as needing clinical guidance, and the eating-disorder rule sending "I hide wrappers" to the team instead of the self-assessment lesson written for it. Gold questions stopping before retrieval went from 2 to 0
Uploads keep the CSV in step Done — a commit rewrites data/kb/knowledge_base.csv from the database as its fourth step. Before that an upload left the file stale, and the next kb-db import offered to delete the uploaded record — the mentor's work disappearing at a moment unconnected to the upload that added it
Uploads without a token Done — the token was removed entirely. Anyone who can reach the server can upload, so it must not be exposed publicly as it is
Handover email strictness Done — a substance check turns away greetings, keysmashes and one-word messages before retrieval, and "no lesson fits" no longer emails. A sixteen-message probe went from eleven emails to one, and that one is the clinical case
Knowledge base on Google Drive Researched only. Drive content is reachable and comes back in the exact shape kb_parser already reads, so no new parser is needed. Nothing is built: the bot has no Google credentials, and it would need a service account on a shared folder before Drive could be a source

What is left

Not code, mostly. The Disciple adapter needs a portal subdomain, an API key, the bot's numeric author id and a webhook secret — and an answer to the one question in the Disciple notes that is nobody's to decide alone: whether the bot should ever reply in public to a post from someone in crisis.

The handover email needs doctortro.com verified with Resend. Until those DNS records exist the only working sender is Resend's sandbox address, which delivers to the account owner and silently drops everyone else — so adding a teammate to DECIPLE_HANDOVER_TO today would look like it worked and would not. Verified working to the account owner's address on 2026-09-27.

And the one piece of code a public deployment needs: nothing guards /kb/preview, /kb/commit or /kb/reload. Anyone who can reach the server can change what members are told, and rebuilding the retriever is the most expensive thing the process can be asked to do. On a laptop that is the deliberate choice recorded in the architecture notes. On a public URL it is the blocker.

A known limitation, measured

The bot cannot tell "the right lesson is missing" from "the right lesson is here". Asked something the courses are near but do not cover, it offers the closest lesson it has, at ordinary confidence.

Measured on 2026-09-17 by deleting each lesson in turn and re-asking its gold question: all 64 stayed above DECIPLE_SCOPE_FLOOR, on a substitute. The floor is one absolute cosine threshold, and the two cases overlap almost exactly — 0.580–0.880 with the lesson present, 0.580–0.777 without it. Raising the floor to 0.75 would catch 91% of the missing-lesson cases and reject 55% of correct answers. The cross-encoder separates them no better. Shown 12 shortlists with the right lesson removed, the model itself declined twice.

The floor still does the job it was built for: a genuinely distant question ("what's the best pizza topping?") scores below it and gets the scope reply. What it cannot detect is missing coverage, and that needs a different mechanism — a second model call asking whether the chosen lesson answers the question, or a reply that offers the lesson as the closest thing found rather than as the answer. Neither is built, and both need their own eval.