build-with-mhonn

Side build

Self-Hosted RAG Document API

Ask questions about your own documents, answered by a model running on your machine. Chunking, embedding, retrieval and generation all run on the host.

The problem

Asking questions about private documents means pasting them into a hosted chat tool, which sends them to someone else's server. Small local models fix that but cannot take a whole document, so something has to decide which passages they see.

What I built

  • Documents are split into overlapping chunks and embedded locally with Ollama.
  • Retrieval is filtered by owner, so one person's chunks never leak into another's answer.
  • Answers stream back as newline-delimited JSON, so cited chunks appear before the first token.
  • Everything except the two model calls is deterministic Python.
  • Includes a browser UI and a guide to running Ollama on your own machine.

Screenshots