MLflow
Every question a member asks and every evaluation run is logged. This page is what is logged, where, and how to read it. Why there are three evaluation tools rather than one is in Architecture, under "Three evaluation tools, each answering a different question".
Opening it
mlflow ui --backend-store-uri sqlite:///mlflow.db
Then http://127.0.0.1:5000. The store is MLFLOW_TRACKING_URI, which
defaults to sqlite:///mlflow.db in the repository root.
Note the port. mkdocs serve also uses 8000 and uvicorn defaults to it, so
running all three at once needs one of them moved.
The experiments
Runs and traces are kept apart by what they answer, so a search for one is not buried in the other. Four are written by something you can run today:
| Experiment | Written by | One entry is |
|---|---|---|
ask |
The bot, every request | One member's question, start to finish |
retrieval-eval |
kb-eval |
One scored run over the gold questions |
answer-eval |
kb-grade |
One scored run over served answers |
trace-judge |
kb-judge |
One LLM reading of the traces |
A fifth, chunking-strategy, is history rather than a live experiment: the
runs behind the chunking comparison in
Architecture, kept so the numbers quoted there can be
checked rather than taken on trust. Nothing writes to it any more.
Reading a request trace
Every question writes one trace to ask, with a span per stage. A trace is
written on every exit, including the ones that stop before retrieval — a
question that never reached the model is exactly the kind you want to find
later.
| Span | What it records |
|---|---|
ask |
The root: the outcome, the reason, whether it emailed |
safety |
What the word list judged — crisis, sensitive or neither |
intent |
Whether the message is a question worth searching for |
confidence |
The cosine score, the floor, and whether it fell back to the whole message |
retrieve |
The four candidates, ranked, with scores |
generate |
What the model was shown and what it chose |
validate |
Which checks passed, and whether the draft was servable |
This is how "why did this member get a handover?" is answered without
re-running anything. The confidence span says whether it was refused as
off-topic and at what score; retrieve says what it had to choose among; and
generate says what it did with them.
Handover emails carry the trace id, so an email and its trace are one click apart. See Handover.
Evaluation runs
kb-eval scores retrieval against the gold questions and logs, per method
— cosine, BM25, RRF fusion and the cross-encoder rerank:
recall_at_1, recall_at_3, mrr, top_score, margin
Parameters go with them: the embedding model, top_k, rrf_k, how many
questions and which file. That is what makes two runs comparable — a recall
number without the embedding model beside it cannot be compared to anything.
It also writes one trace per question, query in and ranked lessons out, per
method.
kb-grade scores answers the bot actually served: accuracy, correct,
wrong_lesson, handed_over, refused, errors, and the two that say whose
fault a wrong answer was — blamed_on_retrieval and blamed_on_model. That
split is the point: a wrong lesson because the right one was never retrieved is
a different problem from a wrong lesson chosen out of a good shortlist.
kb-judge reads the traces with an LLM and logs judged,
skipped_safety, skipped_incomplete, prompt_faults, errors, and a count
per verdict.
Traces are for reading, not just counting
The numbers say whether retrieval is good enough. The traces say what happened to one person. Both matter, and they answer different questions — a run with good recall can still contain a member who got something odd, and only the trace shows you why.
When it is not running
The bot does not require MLflow. If the tracking store is unreachable the traces are filed wherever MLflow's own defaults point, and the member still gets their answer. A surface must not fail to reply because a trace store is down.