Skip to content
AI Agent Campfire

Sherlock Holmes RAG Agent

Sherlock Holmes silhouette with magnifying glass

I wanted to learn more about RAG in AI so I built a Sherlock Holmes RAG Agent, a chatbot that leverages natural language processing (NLP) and machine learning (ML) to answer questions about the Sherlock Holmes canon.

The motivation behind this project was to explore the potential of RAG chatbots in providing accurate and relevant answers to complex questions, while also showcasing the Sherlock Holmes canon in a new and engaging way. Through this project, I aimed to learn more about the capabilities and limitations of RAG chatbots, as well as the challenges and opportunities of integrating NLP and ML with a large corpus of text.

The Sherlock Holmes canon by Arthur Conan Doyle comprises 4 novels and 56 short stories (60 works total), all now in the public domain. I deliberately limited the corpus to one novel and one short-story collection (12 stories) = 13 documents. This is enough to demonstrate a working RAG pipeline, keep embedding cost/latency trivial (all local), and make manual inspection of retrieval quality feasible during evaluation.

# Work Type Source
1 A Study in Scarlet (1887) Novel Project Gutenberg eBook #244
2 The Adventures of Sherlock Holmes (1892) 12 short stories Project Gutenberg eBook #1661

Both texts were downloaded as plain-text UTF-8 from Project Gutenberg’s stable cache URLs:

  • A Study in Scarlet — https://www.gutenberg.org/cache/epub/244/pg244.txt
  • The Adventures of Sherlock Holmes — https://www.gutenberg.org/cache/epub/1661/pg1661.txt
  • 13 documents (1 novel + 12 stories).
  • 147,850 words across the 13 documents that will be indexed.
flowchart TD
    A[Raw text files<br/>data/raw/*.txt] --> B[Ingest<br/>load + normalize]
    B --> C[Chunk<br/>semantic + token-window]
    C --> D[Embed<br/>Ollama nomic-embed-text]
    D --> E[Store<br/>ChromaDB + metadata]
    E --> F[Retrieve<br/>query embed + top-k cosine]
    F --> G{Rerank?<br/>deferred in v1}
    G --> H[Generate<br/>Ollama LLM + cited context]
    H --> I[Cite + return<br/>story + chunk refs]
    J[User query] --> F
    I --> K[User answer + citations]
    K -.-> L[Eval runner<br/>qa_set.json + rubric<br/>Ollama judge]

semantic-first, token-capped, with overlap

Section titled “semantic-first, token-capped, with overlap”

Split on natural boundaries (blank-line paragraph breaks, and the chapter/section markers present in the novel) first, then merge small neighbours and split large ones against a target of ~400 tokens with a 50-token overlap, clamped to a hard range of 150–600 tokens.

This is a hybrid: pure fixed-window token splitting would cut dialogue and deductions mid-sentence (the corpus is heavily dialogue-driven); pure paragraph splitting would produce a long tail of 1-line fragments and a few multi-page monologues that destroy retrieval precision. The hybrid keeps Holmes’s reasoning chains intact while staying inside a model-friendly window.

You are a Sherlock Holmes canon expert answering a reader's question.
You have access ONLY to the retrieved passages below, each labelled
[CHUNK <chunk_id> | <story_title> | <section>]. The corpus you are
drawing from is the novel "A Study in Scarlet" plus the 12 stories of
"The Adventures of Sherlock Holmes" — nothing else.
RULES:
1. Answer using ONLY the information in the provided chunks. If the
question asks about a story, character, or event not covered by any
chunk, say exactly: "I don't have any source material that covers
that." Do NOT use your own knowledge of the Sherlock Holmes stories
beyond what is in the chunks.
2. If the question's premise is wrong (e.g. it mis-attributes a fact to
the wrong story), say so explicitly and name the correct story if a
chunk supports it.
3. If the answer requires information from multiple chunks, combine
them and cite each.
4. If the chunks partially answer the question, say what they support
and explicitly flag what is missing rather than guessing.
5. Cite every factual claim with [CHUNK <chunk_id>]. Place the citation
immediately after the claim, not in a pile at the end.
OUTPUT FORMAT:
- A direct answer in 1–4 sentences.
- A "Sources:" line listing each cited chunk as:
<story_title> — chunk <chunk_index> (<source_file>)
- If you are refusing per rule 1, the answer is the refusal sentence
alone, with an empty Sources line.
RETRIEVED CHUNKS:
[CHUNK 01_a_scandal_in_bohemia#0007 | A Scandal in Bohemia | II]
...chunk text...
[CHUNK 06_the_man_with_the_twisted_lip#0012 | The Man with the Twisted Lip | I]
...chunk text...
(… up to 5 chunks …)
QUESTION: <user query>
  • Python 3.11+
  • Vector DB: ChromaDB (local, persistent, no external infra)
  • Embeddings: Ollama nomic-embed-text (free, local, 768-dim)
  • LLM (generation + judge): Ollama llama3.1 (free, local, no API keys)
  • UI: Streamlit