Skip to content

Setup

Written for Windows PowerShell, which is the development environment. The only commands that are actually Windows-specific are noted with their macOS and Linux equivalents as they come up; everything else — pip, npm, the kb-* commands, uvicorn — is the same everywhere.

Python 3.12 or newer is required. Note that Python 3.12 no longer ships binary installers, so 3.13 is fine and is what you will most likely install.

What is not in the repository

Four things, each left out on purpose, and each needed before the bot answers anything. A clone that has none of them still installs, still passes its tests, and still starts a server — it just has nothing to say.

Missing Why How to get it
Model weights (models/) 216MB, and one file is over GitHub's 100MB hard limit python scripts/fetch_models.py — see Fetching the models
Knowledge base content (data/kb/, data/kb_archive/) It is the mentor's course material, not code. It does not belong in a public repository From whoever owns the content — see step 5
API keys (.env) Secrets Copy-Item .env.example .env and fill in, see step 4
Vector store and database (data/chroma/, data/kb.db) Derived from the two above, so committing them would be committing the content twice Built by the pipeline, step 6

Nothing else is machine-specific. There are no absolute paths anywhere in the repository, and web/.env.local — which holds whichever port the API was started on — is deliberately untracked, so your copy is yours.

1. Clone and branch

git clone <repo-url> deciple_bot
cd deciple_bot
git checkout smit/week3-disciple

2. Create and activate a virtual environment

python -m venv .venv
.\.venv\Scripts\Activate.ps1     # macOS / Linux: source .venv/bin/activate

On Windows, if PowerShell refuses to run the activation script ("running scripts is disabled on this system"), allow local scripts for your user account and try again:

Set-ExecutionPolicy -Scope CurrentUser -ExecutionPolicy RemoteSigned

Confirm the interpreter once the prompt shows (.venv):

python --version

3. Install dependencies

pip install -e ".[dev]"

This installs the RAG stack (LangChain, ChromaDB, sentence-transformers, MLflow, DSPy) plus the dev extra (pytest, MkDocs). It takes a while the first time — sentence-transformers pulls in PyTorch.

Add pip install -e ".[api]" to run the server, which brings in FastAPI and uvicorn. The pipeline and the tests do not need it.

4. Create your .env

Copy-Item .env.example .env     # macOS / Linux: cp .env.example .env

Two values matter to get the pipeline running.

GROQ_API_KEY is required from Week 2 on — the reply layer cannot run without one. Get it at https://console.groq.com/keys. Everything else (parsing, ingest, retrieval, eval) still runs with it empty, and the failure when it is missing is raised at startup rather than on a member's first question.

DECIPLE_EMBEDDING_MODEL ships as the HuggingFace id BAAI/bge-small-en-v1.5, which downloads on first use. Once scripts/fetch_models.py has put the weights in models/, change it to models/bge-small-en-v1.5 — otherwise the pipeline ignores the 216MB you just downloaded and fetches the model again.

DECIPLE_EMBEDDING_MODEL accepts either a HuggingFace repo id, downloaded on first use, or a path to a local folder. Fetching the models locally is the better default either way, and it is the only thing that works on a machine that cannot verify HuggingFace's TLS certificate — which is the case on at least one dev machine here. See Fetching the models.

More matter once the bot is answering members. Each is explained where it is declared in .env.example:

  • RESEND_API_KEY, DECIPLE_MAIL_FROM, DECIPLE_HANDOVER_TO send the handover email. With the key empty the email is composed and written to the log instead of sent, so everything else runs unchanged before mail is set up.
  • DECIPLE_DISCLAIMER is the sentence every reply carries. It ships with a working default, so a deployment that never sets it is still compliant; whoever owns compliance owns the wording.

5. Add knowledge base files

The mentor supplies KB content as batch files. They go here:

data/kb/

These are gitignored; only data/kb/.gitkeep is tracked, to keep the directory in the repo.

6. Run the pipeline

With the KB files in place, four commands in this order:

kb-parse    # parse and report, before anything embeds
kb-db import  # load the files into the database the bot reads
kb-ingest   # embed changed records into Chroma
kb-eval     # score retrieval, log metrics and traces to MLflow

kb-db import is the step that makes the files count. The database named by DECIPLE_DATABASE_URL is what the server reads and what an upload writes to; data/kb/ is where content is imported from and exported to. kb-ingest with no arguments reads the database — pass it a directory to embed files instead, which is how a corpus gets checked before it is imported.

kb-db status says how many records are stored, and kb-db export writes them back out as a CSV.

7. Run the bot

Two processes: the API, and the React app in front of it.

pip install -e ".[api]"
uvicorn deciple_bot.api:app --reload
cd web
npm install
Copy-Item .env.example .env      # or: cp .env.example .env — only if the API is not on :8000
npm run dev

The chat is at http://localhost:5173. The knowledge base upload screen is the same app at http://localhost:5173/admin. It takes no token: anyone who can reach the server can upload. Answering a question needs GROQ_API_KEY; /health and the upload screen do not.

To send one real handover email end to end, once the Resend values are filled in:

python scripts/send_test_handover.py

The unit tests prove the wording and the request are right. What they cannot prove is that your key works, that your sending domain is verified, and that the mail lands in an inbox rather than a spam folder. That is what this is for.

8. Serve the docs

mkdocs serve -a 127.0.0.1:8010

Then open http://127.0.0.1:8010. The port is given explicitly because mkdocs defaults to 8000 and so does the API — running both with the defaults means one of them does not start.

Configuration

All provider choices — the LLM, the embedding model, filesystem paths, and retrieval constants — live in a single config module (src/deciple_bot/config.py), read from environment variables named as in .env.example. Nothing provider-specific is hardcoded anywhere else.

This is not tidiness. The brief requires that the system keeps working if a component is swapped, and that requirement only holds if a swap is a config change rather than a code change. One module is the only place that guarantee can be enforced.

Fetching the models

models/ is gitignored, so a fresh clone has neither model. One command gets both, on any operating system:

python scripts/fetch_models.py

./scripts/fetch_models.ps1 does the same thing and came first, but it reaches for curl.exe and .venv/Scripts/python.exe, so it is Windows only. Use the Python one unless you have a reason not to.

It is safe to re-run — files already present are skipped — and it verifies both models load before it exits.

Why they are not in the repository. They are 216MB together, and bge-small-en-v1.5/model.safetensors alone is 127MB, over GitHub's 100MB per-file hard limit: a push containing it is rejected outright. Git LFS would carry them, at the cost of every clone drawing on a 1GB/month bandwidth quota to fetch bytes that are identical to a public download. So they are fetched rather than committed, and the repository stays under a megabyte.

Neither script uses huggingface_hub: the Hub client wants a cache directory, an offline flag and a revision, and this needs ten files at a URL. The PowerShell one reaches for curl.exe --ssl-no-revoke because Windows' certificate revocation check is what fails behind TLS interception. The Python one uses httpx, which makes no such check, so there is nothing to skip.

The embedding model is required for kb-ingest and every retrieval method. The reranker is required only by the reranked method — the rest of the pipeline runs without it.

Choosing the LLM

DECIPLE_MODEL is a full LiteLLM model id, provider prefix included, so swapping provider is a config change rather than a code change.

The default is groq/openai/gpt-oss-120b — an OpenAI open-weights model served by Groq, which reads oddly and is correct. Groq has retired their Llama chat models: llama-3.3-70b-versatile returns "does not exist or you do not have access to it", llama-3.1-70b-versatile reports as decommissioned, and the only meta-llama entries left on the account are prompt-injection classifiers.

The models on this account that can hold a conversation and have no built-in web search are openai/gpt-oss-120b, openai/gpt-oss-20b and qwen/qwen3.6-27b. Measured against the reply rules as they stood then — chosen id must be one of the candidates, no URL in the sentence, one sentence, no quoting the summary; the cap became three sentences on 2026-09-04 — gpt-oss-120b and qwen3.6-27b both passed every check on every test question. gpt-oss-120b was taken because it is roughly twice as fast and spends fewer tokens per call, which matters against the rate limit below.

groq/compound and groq/compound-mini are excluded deliberately. They work, and they have web search built in. A model that can reach outside the knowledge base breaks the one rule this bot has.

Two things about calling Groq from here

TLS interception. Python ships its own certificate bundle instead of reading the one Windows trusts, so on a machine that intercepts TLS every Groq call fails with CERTIFICATE_VERIFY_FAILED. build_lm calls truststore.inject_into_ssl() first, which makes Python use the operating system trust store — the one that already has the proxy's root certificate in it. It runs unconditionally; on a machine with no interception it changes nothing. This is the same problem that makes the model weights need curl --ssl-no-revoke.

Rate limit. The free tier allows 8000 tokens per minute, and a request costs roughly 1300 — about six questions a minute before Groq returns RateLimitError. Enough for development, not enough to run the reply layer over the whole gold question set in one pass.

Caveat: changing the embedding model invalidates the vector store

Swapping DECIPLE_EMBEDDING_MODEL is not a drop-in change. Every record must be re-ingested, because vectors from a different model are not comparable to the ones already stored. The retrieval threshold must also be re-calibrated: distances from a different model land on a different scale, so a threshold tuned for one model is meaningless for another.

This no longer fails silently: ingest stamps the model name on the collection at creation, and both ingest and retrieval refuse to open a store built with a different model until it is deleted and re-ingested.