Skip to content

MLflow

Every question a member asks and every evaluation run is logged. This page is what is logged, where, and how to read it. Why there are three evaluation tools rather than one is in Architecture, under "Three evaluation tools, each answering a different question".

Opening it

mlflow ui --backend-store-uri sqlite:///mlflow.db

Then http://127.0.0.1:5000. The store is MLFLOW_TRACKING_URI, which defaults to sqlite:///mlflow.db in the repository root.

Note the port. mkdocs serve also uses 8000 and uvicorn defaults to it, so running all three at once needs one of them moved.

The experiments

Runs and traces are kept apart by what they answer, so a search for one is not buried in the other. Four are written by something you can run today:

Experiment Written by One entry is
ask The bot, every request One member's question, start to finish
retrieval-eval kb-eval One scored run over the gold questions
answer-eval kb-grade One scored run over served answers
trace-judge kb-judge One LLM reading of the traces

A fifth, chunking-strategy, is history rather than a live experiment: the runs behind the chunking comparison in Architecture, kept so the numbers quoted there can be checked rather than taken on trust. Nothing writes to it any more.

Reading a request trace

Every question writes one trace to ask, with a span per stage. A trace is written on every exit, including the ones that stop before retrieval — a question that never reached the model is exactly the kind you want to find later.

Span What it records
ask The root: the outcome, the reason, whether it emailed
safety What the word list judged — crisis, sensitive or neither
intent Whether the message is a question worth searching for
confidence The cosine score, the floor, and whether it fell back to the whole message
retrieve The four candidates, ranked, with scores
generate What the model was shown and what it chose
validate Which checks passed, and whether the draft was servable

This is how "why did this member get a handover?" is answered without re-running anything. The confidence span says whether it was refused as off-topic and at what score; retrieve says what it had to choose among; and generate says what it did with them.

Handover emails carry the trace id, so an email and its trace are one click apart. See Handover.

Evaluation runs

kb-eval scores retrieval against the gold questions and logs, per method — cosine, BM25, RRF fusion and the cross-encoder rerank:

recall_at_1, recall_at_3, mrr, top_score, margin

Parameters go with them: the embedding model, top_k, rrf_k, how many questions and which file. That is what makes two runs comparable — a recall number without the embedding model beside it cannot be compared to anything. It also writes one trace per question, query in and ranked lessons out, per method.

kb-grade scores answers the bot actually served: accuracy, correct, wrong_lesson, handed_over, refused, errors, and the two that say whose fault a wrong answer was — blamed_on_retrieval and blamed_on_model. That split is the point: a wrong lesson because the right one was never retrieved is a different problem from a wrong lesson chosen out of a good shortlist.

kb-judge reads the traces with an LLM and logs judged, skipped_safety, skipped_incomplete, prompt_faults, errors, and a count per verdict.

Traces are for reading, not just counting

The numbers say whether retrieval is good enough. The traces say what happened to one person. Both matter, and they answer different questions — a run with good recall can still contain a member who got something odd, and only the trace shows you why.

When it is not running

The bot does not require MLflow. If the tracking store is unreachable the traces are filed wherever MLflow's own defaults point, and the member still gets their answer. A surface must not fail to reply because a trace store is down.