Side build
Self-Hosted RAG Document API
Ask questions about your own documents, answered by a model running on your machine. Chunking, embedding, retrieval and generation all run on the host.
The problem
Asking questions about private documents means pasting them into a hosted chat tool, which sends them to someone else's server. Small local models fix that but cannot take a whole document, so something has to decide which passages they see.
What I built
- Documents are split into overlapping chunks and embedded locally with Ollama.
- Retrieval is filtered by owner, so one person's chunks never leak into another's answer.
- Answers stream back as newline-delimited JSON, so cited chunks appear before the first token.
- Everything except the two model calls is deterministic Python.
- Includes a browser UI and a guide to running Ollama on your own machine.
Screenshots
- 01Screenshot coming soon
Swagger UI - 02Screenshot coming soon
Question with cited chunks - 03Screenshot coming soon
Embeddings