Skip to content
AI Agent Campfire

What I Learned

Building a RAG agent quickly made one thing clear: getting an LLM to answer questions is not the difficult part.

The more interesting challenge is making sure the model receives the right information, in the right form, at the right time.

For this experiment, I focused on three areas that had the biggest impact on the quality of the system:

  • Chunking strategies
  • Retrieval and evaluation
  • Iterative experimentation

Rather than treating these as independent components, I looked at them as parts of the same feedback loop:

Documents → Chunking → Embeddings → Retrieval → Context → Answer → Evaluation → Iteration

1. Chunking Is More Important Than It Looks

Section titled “1. Chunking Is More Important Than It Looks”

The first assumption I wanted to test was that splitting documents into reasonably sized chunks would be enough.

It wasn’t.

Different chunking strategies can produce very different retrieval results, even when the underlying documents and embedding model remain unchanged.

I experimented with factors such as:

  • Chunk size
  • Chunk overlap
  • Splitting by paragraphs versus fixed token/character sizes
  • Preserving semantic boundaries
  • The amount of surrounding context included with a chunk

The interesting finding was that chunking is not simply a preprocessing decision.

It directly affects what the retrieval system can find.

A chunk that is too small may contain insufficient context to answer a question. A chunk that is too large may contain too much unrelated information and make retrieval less precise.

The evaluation results reinforced this point. Several questions had the correct information somewhere in the corpus, but the relevant chunk did not make it into the top-5 retrieved results.

That means a system can have the right information available and still produce the wrong answer simply because the retrieval layer failed to surface it.

Lesson: Chunking is part of retrieval design. Before changing the LLM or embedding model, it is worth asking whether the information is being divided in a way that makes sense for the questions the system needs to answer.

2. Evals Turned Guesswork Into Experimentation

Section titled “2. Evals Turned Guesswork Into Experimentation”

The second challenge was more fundamental:

How do I know whether one approach is actually better than another?

Without evaluation, it is surprisingly easy to look at a few successful answers and conclude that the RAG system is working.

That isn’t enough.

The first full end-to-end evaluation contained 30 questions across four categories: factual single-hop, multi-hop, cross-story consistency, and out-of-scope questions.

Factual single-hop: The answer is sitting in one place. You ask a question, the system grabs one relevant passage, and the answer is right there. Example: “Who is Sherlock Holmes’ roommate?” — one chunk from one story has the answer. Multi-hop: The answer needs pieces from two or more places stitched together. No single passage has the full answer — the system has to retrieve multiple chunks and combine them to reason it out. Example: “How did Holmes and Watson first meet, and what case did they solve together right after?” — that likely spans two different passages or even two stories. Cross-story consistency: Similar to multi-hop, but specifically about comparing or reconciling facts that appear in different stories, especially where the answer requires checking multiple stories at once to be accurate or to catch a contradiction. Example: “Does Watson’s war wound show up in the same place across different stories?” — you have to pull from several stories and compare, not just combine. Out-of-scope: The question is about something that simply isn’t in your corpus (like a story you didn’t include, or a made-up fan theory) — so a good system should say “I don’t know / not in my source material” instead of making something up. It’s a test of honesty, not retrieval skill.

The headline result was encouraging:

0% hallucination rate — 0 out of 30 answers contained a confidently fabricated specific detail.

However, the overall correctness was only 57% (17/30).

Evaluation area Correct Partial Wrong
Factual single-hop 5/9 3/9 1/9
Multi-hop 2/6 2/6 2/6
Cross-story consistency 0/5 2/5 3/5
Out-of-scope 10/10 0/10 0/10
Overall 17/30 7/30 6/30

This was actually more useful than simply seeing a high accuracy number.

The evaluation showed where the system was failing and, more importantly, why.

Five of the six wrong answers were primarily associated with retrieval misses.

The relevant information existed in the corpus, but the chunk containing the key fact was not included in the top-5 retrieved results.

This produced an important distinction:

The system wasn’t necessarily unable to answer the question. It often wasn’t given the information required to answer it.

For example, one question asked who killed Charles McCarthy and what weapon was used. The system identified the correct killer but gave the wrong weapon because the chunk describing the actual weapon—a jagged stone—was not retrieved. A different chunk containing a pen-knife reference was retrieved instead, leading the model to conflate two pieces of information.

That was a particularly useful failure because it demonstrated how a retrieval problem can eventually look like a generation problem.

Lesson: When a RAG answer is wrong, don’t immediately blame the LLM. First ask: Was the evidence required to answer the question actually present in the retrieved context?

3. Multi-Hop Questions Exposed the Limits of Top-K Retrieval

Section titled “3. Multi-Hop Questions Exposed the Limits of Top-K Retrieval”

The most interesting weakness appeared in the multi-hop and cross-story questions.

The system used k=5 retrieval, but some questions required evidence from multiple stories.

The cross-story consistency tier scored 0/5 correct.

The reason became fairly clear during analysis: with roughly 370 chunks across 13 documents, retrieving only five chunks often meant that the system received information from one or two stories while missing the other stories required for comparison.

This created an important failure pattern.

A question might ask:

Compare what happened in Story A and Story B.

But the retrieval layer effectively answered:

Here is some relevant information about Story A.

The generation model then had no reliable basis for performing the comparison.

This suggests that increasing k for questions that explicitly mention multiple stories could be a relatively inexpensive next experiment.

A more advanced approach would be iterative retrieval:

Retrieve → identify missing entities/stories → retrieve again → combine evidence → answer

That approach was outside the scope of this first experiment, but the evaluation made the potential value of it much clearer.

Lesson: Retrieval parameters should reflect the complexity of the question. A single fixed k may not be appropriate for both simple factual questions and multi-document reasoning.

4. Not Every Failure Was a Retrieval Problem

Section titled “4. Not Every Failure Was a Retrieval Problem”

One of the useful aspects of running evaluations was discovering that retrieval did not explain everything.

One multi-hop question deliberately contained a false premise: it assumed that two stories contained women who disguised themselves as men.

The correct response was to challenge the premise.

Instead, the model attempted to satisfy the requested count and produced unsupported examples.

This was primarily a generation/prompt-following failure, rather than a retrieval failure.

That distinction matters.

The evaluation therefore helped separate the system’s problems into different categories:

Retrieval failure → The required evidence was not retrieved.

Generation failure → The evidence was available, but the model interpreted or followed the instruction incorrectly.

Evaluation failure → The evaluator itself misclassified the result.

This categorization makes future improvements much more targeted.

One unexpected finding came from the evaluation process itself.

One question was incorrectly classified as a successful refusal by the LLM judge.

The system said it did not have enough source material, and the judge marked that as correct. However, manual inspection showed that the answer was actually in scope and the system had simply failed to retrieve the relevant chunks.

This is an important reminder that an LLM-as-judge is not automatically ground truth.

The evaluator can have its own blind spots, particularly when it cannot independently inspect the full corpus.

For this experiment, that means evaluation should include some level of manual calibration, especially for difficult multi-hop and cross-document questions.

Lesson: An evaluation system is itself part of the system being designed. A score is only as trustworthy as the evaluation methodology behind it.

The most valuable outcome of the evaluation wasn’t the 57% correctness score or even the 0% hallucination rate.

It was the ability to identify the next engineering questions.

The results suggest several concrete experiments:

  • Increase k for multi-story questions.
  • Compare different chunk sizes and boundaries.
  • Experiment with semantic versus fixed-size chunking.
  • Investigate hybrid retrieval.
  • Add reranking improvements.
  • Explore iterative retrieval for multi-hop questions.
  • Strengthen the prompt rule for challenging false premises.
  • Improve judge calibration for false refusals.
  • Expand the evaluation dataset before drawing broader conclusions.

This creates a much more useful development cycle:

Hypothesis → Change one variable → Run evaluation → Inspect failures → Identify root cause → Iterate

My biggest takeaway from this experiment is that RAG quality is not determined by a single component.

It is an interaction between the data, chunking strategy, retrieval method, context construction, generation, and evaluation.

More importantly, improving the system is less about finding a perfect configuration and more about creating a reliable way to measure, experiment, and learn.

The first evaluation changed my perspective in an interesting way.

I initially expected the main challenge to be preventing the LLM from hallucinating.

Instead, the strongest result was that hallucination was not the dominant problem.

The bigger problem was getting the right evidence into the context in the first place.

That led to a simple conclusion:

Don’t start by asking “Which model should I use?” Start by asking “How will I know that my system is getting better?”

Once that question is answered, chunking strategies, retrieval methods, prompts, embedding models, and other changes become experiments that can be measured rather than guesses that happen to work on a few examples.

This experiment is intentionally simple and is primarily a learning project. The results are therefore not intended to represent production-level RAG performance.

The next experiments I would explore are:

  • More systematic chunking strategies
  • Dynamic k based on query complexity
  • Hybrid search combining semantic and keyword retrieval
  • Reranking retrieved documents
  • Metadata-aware retrieval
  • Query rewriting
  • Iterative retrieval for multi-hop questions
  • More comprehensive evaluation datasets
  • Retrieval-specific metrics
  • Evaluation of groundedness and answer relevance
  • Better calibration of LLM-as-judge evaluations
  • Comparing different embedding models

The goal is not to build the “perfect” RAG system.

The goal is to understand why a RAG system succeeds or fails—and develop a repeatable way to improve it.