Sherlock Holmes RAG Agent

I wanted to learn more about RAG in AI so I built a Sherlock Holmes RAG Agent, a chatbot that leverages natural language processing (NLP) and machine learning (ML) to answer questions about the Sherlock Holmes canon.
Why I built Sherlock Holmes Rag Agent
Section titled “Why I built Sherlock Holmes Rag Agent”The motivation behind this project was to explore the potential of RAG chatbots in providing accurate and relevant answers to complex questions, while also showcasing the Sherlock Holmes canon in a new and engaging way. Through this project, I aimed to learn more about the capabilities and limitations of RAG chatbots, as well as the challenges and opportunities of integrating NLP and ML with a large corpus of text.
The Approch
Section titled “The Approch”The Sherlock Holmes canon by Arthur Conan Doyle comprises 4 novels and 56 short stories (60 works total), all now in the public domain. I deliberately limited the corpus to one novel and one short-story collection (12 stories) = 13 documents. This is enough to demonstrate a working RAG pipeline, keep embedding cost/latency trivial (all local), and make manual inspection of retrieval quality feasible during evaluation.
Included
Section titled “Included”| # | Work | Type | Source |
|---|---|---|---|
| 1 | A Study in Scarlet (1887) | Novel | Project Gutenberg eBook #244 |
| 2 | The Adventures of Sherlock Holmes (1892) | 12 short stories | Project Gutenberg eBook #1661 |
Source
Section titled “Source”Both texts were downloaded as plain-text UTF-8 from Project Gutenberg’s stable cache URLs:
- A Study in Scarlet —
https://www.gutenberg.org/cache/epub/244/pg244.txt - The Adventures of Sherlock Holmes —
https://www.gutenberg.org/cache/epub/1661/pg1661.txt
Corpus totals
Section titled “Corpus totals”- 13 documents (1 novel + 12 stories).
- 147,850 words across the 13 documents that will be indexed.
Architecture
Section titled “Architecture”Pipeline Overview
Section titled “Pipeline Overview”flowchart TD
A[Raw text files<br/>data/raw/*.txt] --> B[Ingest<br/>load + normalize]
B --> C[Chunk<br/>semantic + token-window]
C --> D[Embed<br/>Ollama nomic-embed-text]
D --> E[Store<br/>ChromaDB + metadata]
E --> F[Retrieve<br/>query embed + top-k cosine]
F --> G{Rerank?<br/>deferred in v1}
G --> H[Generate<br/>Ollama LLM + cited context]
H --> I[Cite + return<br/>story + chunk refs]
J[User query] --> F
I --> K[User answer + citations]
K -.-> L[Eval runner<br/>qa_set.json + rubric<br/>Ollama judge]
Chunking Strategy
Section titled “Chunking Strategy”semantic-first, token-capped, with overlap
Section titled “semantic-first, token-capped, with overlap”Split on natural boundaries (blank-line paragraph breaks, and the chapter/section markers present in the novel) first, then merge small neighbours and split large ones against a target of ~400 tokens with a 50-token overlap, clamped to a hard range of 150–600 tokens.
This is a hybrid: pure fixed-window token splitting would cut dialogue and deductions mid-sentence (the corpus is heavily dialogue-driven); pure paragraph splitting would produce a long tail of 1-line fragments and a few multi-page monologues that destroy retrieval precision. The hybrid keeps Holmes’s reasoning chains intact while staying inside a model-friendly window.
Prompt template
Section titled “Prompt template”You are a Sherlock Holmes canon expert answering a reader's question.You have access ONLY to the retrieved passages below, each labelled[CHUNK <chunk_id> | <story_title> | <section>]. The corpus you aredrawing from is the novel "A Study in Scarlet" plus the 12 stories of"The Adventures of Sherlock Holmes" — nothing else.
RULES:1. Answer using ONLY the information in the provided chunks. If the question asks about a story, character, or event not covered by any chunk, say exactly: "I don't have any source material that covers that." Do NOT use your own knowledge of the Sherlock Holmes stories beyond what is in the chunks.2. If the question's premise is wrong (e.g. it mis-attributes a fact to the wrong story), say so explicitly and name the correct story if a chunk supports it.3. If the answer requires information from multiple chunks, combine them and cite each.4. If the chunks partially answer the question, say what they support and explicitly flag what is missing rather than guessing.5. Cite every factual claim with [CHUNK <chunk_id>]. Place the citation immediately after the claim, not in a pile at the end.
OUTPUT FORMAT:- A direct answer in 1–4 sentences.- A "Sources:" line listing each cited chunk as: <story_title> — chunk <chunk_index> (<source_file>)- If you are refusing per rule 1, the answer is the refusal sentence alone, with an empty Sources line.
RETRIEVED CHUNKS:[CHUNK 01_a_scandal_in_bohemia#0007 | A Scandal in Bohemia | II]...chunk text...[CHUNK 06_the_man_with_the_twisted_lip#0012 | The Man with the Twisted Lip | I]...chunk text...(… up to 5 chunks …)
QUESTION: <user query>Tech Stack
Section titled “Tech Stack”- Python 3.11+
- Vector DB: ChromaDB (local, persistent, no external infra)
- Embeddings: Ollama
nomic-embed-text(free, local, 768-dim) - LLM (generation + judge): Ollama
llama3.1(free, local, no API keys) - UI: Streamlit
Sherlock’s Rag Agent: Screen shot
Section titled “Sherlock’s Rag Agent: Screen shot”