Setup
Written for Windows PowerShell, which is the development environment. The
only commands that are actually Windows-specific are noted with their macOS and
Linux equivalents as they come up; everything else — pip, npm, the kb-*
commands, uvicorn — is the same everywhere.
Python 3.12 or newer is required. Note that Python 3.12 no longer ships binary installers, so 3.13 is fine and is what you will most likely install.
What is not in the repository
Four things, each left out on purpose, and each needed before the bot answers anything. A clone that has none of them still installs, still passes its tests, and still starts a server — it just has nothing to say.
| Missing | Why | How to get it |
|---|---|---|
Model weights (models/) |
216MB, and one file is over GitHub's 100MB hard limit | python scripts/fetch_models.py — see Fetching the models |
Knowledge base content (data/kb/, data/kb_archive/) |
It is the mentor's course material, not code. It does not belong in a public repository | From whoever owns the content — see step 5 |
API keys (.env) |
Secrets | Copy-Item .env.example .env and fill in, see step 4 |
Vector store and database (data/chroma/, data/kb.db) |
Derived from the two above, so committing them would be committing the content twice | Built by the pipeline, step 6 |
Nothing else is machine-specific. There are no absolute paths anywhere in the
repository, and web/.env.local — which holds whichever port the API was
started on — is deliberately untracked, so your copy is yours.
1. Clone and branch
git clone <repo-url> deciple_bot
cd deciple_bot
git checkout smit/week3-disciple
2. Create and activate a virtual environment
python -m venv .venv
.\.venv\Scripts\Activate.ps1 # macOS / Linux: source .venv/bin/activate
On Windows, if PowerShell refuses to run the activation script ("running scripts is disabled on this system"), allow local scripts for your user account and try again:
Set-ExecutionPolicy -Scope CurrentUser -ExecutionPolicy RemoteSigned
Confirm the interpreter once the prompt shows (.venv):
python --version
3. Install dependencies
pip install -e ".[dev]"
This installs the RAG stack (LangChain, ChromaDB, sentence-transformers,
MLflow, DSPy) plus the dev extra (pytest, MkDocs). It takes a while the
first time — sentence-transformers pulls in PyTorch.
Add pip install -e ".[api]" to run the server, which brings in FastAPI and
uvicorn. The pipeline and the tests do not need it.
4. Create your .env
Copy-Item .env.example .env # macOS / Linux: cp .env.example .env
Two values matter to get the pipeline running.
GROQ_API_KEY is required from Week 2 on — the reply layer cannot run
without one. Get it at https://console.groq.com/keys. Everything else
(parsing, ingest, retrieval, eval) still runs with it empty, and the failure
when it is missing is raised at startup rather than on a member's first
question.
DECIPLE_EMBEDDING_MODEL ships as the HuggingFace id
BAAI/bge-small-en-v1.5, which downloads on first use. Once
scripts/fetch_models.py has put the weights in models/, change it to
models/bge-small-en-v1.5 — otherwise the pipeline ignores the 216MB you just
downloaded and fetches the model again.
DECIPLE_EMBEDDING_MODEL accepts either a HuggingFace repo id, downloaded on
first use, or a path to a local folder. Fetching the models locally is the
better default either way, and it is the only thing that works on a machine
that cannot verify HuggingFace's TLS certificate — which is the case on at
least one dev machine here. See Fetching the models.
More matter once the bot is answering members. Each is explained where it is
declared in .env.example:
RESEND_API_KEY,DECIPLE_MAIL_FROM,DECIPLE_HANDOVER_TOsend the handover email. With the key empty the email is composed and written to the log instead of sent, so everything else runs unchanged before mail is set up.DECIPLE_DISCLAIMERis the sentence every reply carries. It ships with a working default, so a deployment that never sets it is still compliant; whoever owns compliance owns the wording.
5. Add knowledge base files
The mentor supplies KB content as batch files. They go here:
data/kb/
These are gitignored; only data/kb/.gitkeep is tracked, to keep the directory
in the repo.
6. Run the pipeline
With the KB files in place, four commands in this order:
kb-parse # parse and report, before anything embeds
kb-db import # load the files into the database the bot reads
kb-ingest # embed changed records into Chroma
kb-eval # score retrieval, log metrics and traces to MLflow
kb-db import is the step that makes the files count. The database named by
DECIPLE_DATABASE_URL is what the server reads and what an upload writes to;
data/kb/ is where content is imported from and exported to. kb-ingest with
no arguments reads the database — pass it a directory to embed files instead,
which is how a corpus gets checked before it is imported.
kb-db status says how many records are stored, and kb-db export writes them
back out as a CSV.
7. Run the bot
Two processes: the API, and the React app in front of it.
pip install -e ".[api]"
uvicorn deciple_bot.api:app --reload
cd web
npm install
Copy-Item .env.example .env # or: cp .env.example .env — only if the API is not on :8000
npm run dev
The chat is at http://localhost:5173. The knowledge base upload screen is the
same app at http://localhost:5173/admin. It takes no token: anyone who
can reach the server can upload. Answering a question needs GROQ_API_KEY; /health and
the upload screen do not.
To send one real handover email end to end, once the Resend values are filled in:
python scripts/send_test_handover.py
The unit tests prove the wording and the request are right. What they cannot prove is that your key works, that your sending domain is verified, and that the mail lands in an inbox rather than a spam folder. That is what this is for.
8. Serve the docs
mkdocs serve -a 127.0.0.1:8010
Then open http://127.0.0.1:8010. The port is given explicitly because mkdocs defaults to 8000 and so does the API — running both with the defaults means one of them does not start.
Configuration
All provider choices — the LLM, the embedding model, filesystem paths, and
retrieval constants — live in a single config module
(src/deciple_bot/config.py), read from environment variables named as in
.env.example. Nothing provider-specific is hardcoded anywhere else.
This is not tidiness. The brief requires that the system keeps working if a component is swapped, and that requirement only holds if a swap is a config change rather than a code change. One module is the only place that guarantee can be enforced.
Fetching the models
models/ is gitignored, so a fresh clone has neither model. One command gets
both, on any operating system:
python scripts/fetch_models.py
./scripts/fetch_models.ps1 does the same thing and came first, but it reaches
for curl.exe and .venv/Scripts/python.exe, so it is Windows only. Use the
Python one unless you have a reason not to.
It is safe to re-run — files already present are skipped — and it verifies both models load before it exits.
Why they are not in the repository. They are 216MB together, and
bge-small-en-v1.5/model.safetensors alone is 127MB, over GitHub's 100MB
per-file hard limit: a push containing it is rejected outright. Git LFS would
carry them, at the cost of every clone drawing on a 1GB/month bandwidth quota
to fetch bytes that are identical to a public download. So they are fetched
rather than committed, and the repository stays under a megabyte.
Neither script uses huggingface_hub: the Hub client wants a cache directory,
an offline flag and a revision, and this needs ten files at a URL. The
PowerShell one reaches for curl.exe --ssl-no-revoke because Windows'
certificate revocation check is what fails behind TLS interception. The Python
one uses httpx, which makes no such check, so there is nothing to skip.
The embedding model is required for kb-ingest and every retrieval method. The
reranker is required only by the reranked method — the rest of the pipeline
runs without it.
Choosing the LLM
DECIPLE_MODEL is a full LiteLLM model id, provider prefix included, so
swapping provider is a config change rather than a code change.
The default is groq/openai/gpt-oss-120b — an OpenAI open-weights model served
by Groq, which reads oddly and is correct. Groq has retired their Llama chat
models: llama-3.3-70b-versatile returns "does not exist or you do not have
access to it", llama-3.1-70b-versatile reports as decommissioned, and the
only meta-llama entries left on the account are prompt-injection classifiers.
The models on this account that can hold a conversation and have no built-in
web search are openai/gpt-oss-120b, openai/gpt-oss-20b and
qwen/qwen3.6-27b. Measured against the reply rules as they stood then —
chosen id must be one of the candidates, no URL in the sentence, one sentence,
no quoting the summary; the cap became three sentences on 2026-09-04 —
gpt-oss-120b and qwen3.6-27b both passed every check on every test question.
gpt-oss-120b was taken because it is roughly twice as fast and spends fewer
tokens per call, which matters against the rate limit below.
groq/compound and groq/compound-mini are excluded deliberately. They work,
and they have web search built in. A model that can reach outside the knowledge
base breaks the one rule this bot has.
Two things about calling Groq from here
TLS interception. Python ships its own certificate bundle instead of
reading the one Windows trusts, so on a machine that intercepts TLS every Groq
call fails with CERTIFICATE_VERIFY_FAILED. build_lm calls
truststore.inject_into_ssl() first, which makes Python use the operating
system trust store — the one that already has the proxy's root certificate in
it. It runs unconditionally; on a machine with no interception it changes
nothing. This is the same problem that makes the model weights need
curl --ssl-no-revoke.
Rate limit. The free tier allows 8000 tokens per minute, and a request
costs roughly 1300 — about six questions a minute before Groq returns
RateLimitError. Enough for development, not enough to run the reply layer
over the whole gold question set in one pass.
Caveat: changing the embedding model invalidates the vector store
Swapping DECIPLE_EMBEDDING_MODEL is not a drop-in change. Every record must
be re-ingested, because vectors from a different model are not comparable to
the ones already stored. The retrieval threshold must also be re-calibrated:
distances from a different model land on a different scale, so a threshold
tuned for one model is meaningless for another.
This no longer fails silently: ingest stamps the model name on the collection at creation, and both ingest and retrieval refuse to open a store built with a different model until it is deleted and re-ingested.