Skip to content
AI Agent Campfire

Case Studies

How I approach problems, design solutions and measure outcomes across AI, product and technology.

Case studies showing how I approach complex problems across product, UX, technology and AI — focusing on the thinking, decisions and outcomes behind the work, not just the technology.


JobScout — Building an Agentic Job-Search Assistant

Section titled “JobScout — Building an Agentic Job-Search Assistant”

Problem: Job seekers are still searching, matching and applying manually while employers use AI to screen them. The same repetitive grind — scan, read, tweak, write, hope — multiplied across 20–30 applications.

Approach: An agentic AI job-search assistant that automates searching, scoring and tailoring resumes with human-in-the-loop approval. It orchestrates multiple LLM-powered modules — resume parsing, semantic matching, tailored content generation and cover-letter writing — in a repeatable scouting cycle.

Architecture: 7-step pipeline (parse → search → dedup → score → filter → tailor → package) orchestrated via a Claude Agent SDK harness with an MCP server that auto-discovers tool modules. Hybrid scoring: 0.6 × cosine similarity + 0.4 × skill overlap using local Ollama embeddings (nomic-embed-text, 768-dim) against a persistent ChromaDB vector store.

Stack: Python 3.11+ · Claude Agent SDK · ChromaDB · Ollama (nomic-embed-text) · Pydantic · Langfuse (telemetry) · 165+ tests with pytest

Key design decision: The agent stops at producing ready-to-review files — it does not submit applications. That boundary was a deliberate design choice, not a missing feature. Human judgment stays in the loop for submission and tracking.


Sherlock Holmes RAG Agent — Retrieval Is the Real Experiment

Section titled “Sherlock Holmes RAG Agent — Retrieval Is the Real Experiment”

Problem: Building a RAG chatbot is easy. Knowing whether it actually works is hard. I wanted to understand what makes a RAG system succeed or fail — not just ship a demo.

Approach: Built a RAG agent over the Sherlock Holmes canon (13 documents, ~148K words) with a semantic-first chunking strategy (~400 tokens, 50-token overlap), local embeddings and a strict prompt template enforcing citation-only answers.

Evaluation: Designed a 30-question test set across four categories — factual single-hop, multi-hop, cross-story consistency and out-of-scope. Used an LLM-as-judge with manual calibration.

Results: 0% hallucination rate (0/30 fabricated details). 57% correctness (17/30). The dominant failure mode was retrieval, not generation — 5 of 6 wrong answers were caused by missing context, not model errors. Cross-story consistency scored 0/5, exposing the limits of fixed top-k retrieval for multi-document reasoning.

Key insight: “Don’t start by asking ‘Which model should I use?’ Start by asking ‘How will I know that my system is getting better?’” — evaluation methodology matters more than model selection.


Every case study follows the same shape:

Problem
↓
Context & research
↓
Opportunity
↓
Options explored
↓
Solution
↓
Design decisions
↓
Implementation
↓
Evaluation
↓
Outcome & learning

The evaluation step is not optional. A system that hasn’t been measured is a prototype, not a solution.