Skip to content

Updating the knowledge base

Four ways content reaches members, and when each one is the right one. The reasoning behind the design is in Architecture, under "The knowledge base is updated by upload, not by a developer".

Which route you want

You have Use Server restart
A file of new or corrected content, and a running bot The upload screen at /admin No
Edits made directly in knowledge_base.csv kb-db import, kb-ingest, kb-reload No
A new batch of the mentor's markdown kb-csv, then the row above No
Nothing but a database, and a lost CSV kb-db export n/a

The upload screen

The one that needs no command line. /admin on the same app as the chat.

  1. Choose a file. .md, .txt or .csv, up to 2MB. The shape is read from the content rather than the extension, so all five of the parser's record shapes work whatever the file is called.
  2. Click "Check this file". This is POST /kb/preview. It reports what the file would add, change and leave alone. It writes nothing and embeds nothing, so it is safe to click on anything.
  3. Click "Add to the knowledge base". This is POST /kb/commit, and it is where everything happens.

A commit does four things before the button stops saying "Adding…":

  • writes the records to the database,
  • embeds the new and changed ones into Chroma,
  • rebuilds the retriever the running server answers from,
  • rewrites data/kb/knowledge_base.csv to match the database.

So when the screen shows the new count, the bot is already answering with the new content. You do not run kb-reload after an upload.

A commit adds and updates. It never removes. Uploading a file holding half the knowledge base does not delete the other half — it reports every row as unchanged and writes nothing. Removing content is a CSV edit, below.

If the CSV rewrite fails — the file open in Excel, a permissions problem — the upload still succeeds, because members can already see it, and the screen says to run kb-db export before the next import.

Editing the CSV

data/kb/knowledge_base.csv is one row per record and opens in Excel. This is the route for corrections and for deletions.

kb-db import   # make the database match the file
kb-ingest      # embed what changed, delete what went away
kb-reload      # tell a running server to re-read the store

npm run ingest from the repository root is all three in one command.

kb-db import replaces. The database ends up holding exactly what the directory holds, so a row you deleted is deleted. That is deliberate — an upsert would silently keep rows a maintainer had removed — but it means an import that would lose records stops and asks:

error: importing data\kb would remove 5 record(s) already in the database.
Re-run with --force if that is what you want.

Read that as a question, not a failure. If you meant to delete those records, --force. If you did not, something is out of date and forcing it would throw away content.

kb-reload is not optional. The retriever loads the whole collection into memory when it is built, so without it the store holds the new record and the running bot cannot find it — the store says 67 records, /health says 66. With no server running it says so and stops, which is the normal case when rebuilding on a laptop.

A new markdown batch

kb-csv reads the mentor's original markdown from data/kb_archive/ and writes data/kb/knowledge_base.csv. Run it when a batch arrives, then kb-db import to make it the records the bot answers from.

It refuses to overwrite an existing CSV without --force, because hand edits in the CSV are not in the markdown. --audit prints the per-file comparison of Q: boundaries against dashed separator lines.

Checking it worked

kb-db status                  # how many records are stored
curl http://127.0.0.1:8000/health   # what the running bot is serving

Those two disagreeing is the symptom kb-reload fixes.

The stronger check is to ask a question whose lesson you just changed. Counts prove a command ran; an answer proves the member sees it.

Worth knowing when you test by deleting: a missing lesson does not produce "I don't have that". The bot offers the closest lesson it still has, at normal confidence. Deleting a lesson and watching the reply change is evidence the ingest worked — it is not evidence the bot noticed anything was missing.

Keeping the file and the database in step

The database is what the server reads. data/kb/ is where content is imported from and exported to. They are kept in step from both ends: kb-db import makes the database match the file, and a commit rewrites the file to match the database.

kb-db export writes the database out as a CSV — the repair when the file is lost, or when the rewrite after an upload failed.