Learning
Experiments, evaluation results and technical notes from building AI and agentic systems.
Learning by building
Section titled “Learning by building”I learn by building real systems and measuring whether they work. Every experiment includes an evaluation step — because a system that hasn’t been measured is a prototype, not a solution.
Current areas
Section titled “Current areas”RAG Systems & Evaluation
Built a RAG agent over the Sherlock Holmes canon with ChromaDB, Ollama embeddings and a 30-question evaluation harness. 0% hallucination rate, 57% correctness — retrieval was the dominant failure mode.
Chunking & Retrieval Strategy
Semantic-first, token-capped chunking (~400 tokens, 50-token overlap) and why chunking is part of retrieval design, not preprocessing.
Agent Harness & Tool Orchestration
Building an agent harness with the Claude Agent SDK and an MCP server that auto-discovers tool modules and exposes them as callable tools.
LLM Observability
Instrumenting every LLM call with Langfuse for token usage, cost tracking and latency monitoring across pipeline stages.
Case Studies
How I approach problems — from problem framing through design decisions to evaluation and outcomes.
Projects
Working prototypes and systems I have built across AI agents, RAG and product engineering.
How I learn
Section titled “How I learn”I combine:
Understand → Experiment → Build → Evaluate → Document
Rather than treating learning as a collection of courses, I use projects and experiments to turn concepts into practical knowledge. Every project gets written up — the writing is where the thinking gets sharp.