Skip to content
AI Agent Campfire

Financial Market Brief Generator

An AI system that generates audited daily market briefs from SEC filings and news — with a Researcher agent, a Critic agent, and a deterministic revision loop

Tracking specific sectors of the financial markets is a daily grind — scanning SEC filings, reading earnings reports, monitoring news feeds, then synthesizing everything into a coherent summary. Most existing tools give you everything; what you need is a focused brief on the sectors you actually care about, with evidence you can trace back to the source.

The harder problem: how do you trust that an AI-generated brief is grounded in real data and not hallucinated? Generating text is easy — ensuring every claim is backed by a cited source, reviewed by a second agent, and revised when it fails review is the actual engineering challenge.


The system is built in two phases — a deterministic data pipeline that produces evidence bundles, and an agent layer (Researcher + Critic) that turns those bundles into audited briefs through a revision loop.

A deterministic Python pipeline that turns a GICS sector into a versioned evidence bundle (JSON) of new SEC filings and news items since the last run. No LLM code — pure data plumbing.

7-step pipeline per sector:

  1. Resolve companies — weekly-cached company list per sector (priority tickers first, then EDGAR full-text search by SIC code)
  2. Load state — per-sector JSON tracking last run timestamp, seen SEC accession numbers, and seen news URL hashes. Computes the since cutoff (last run, or now − 7 days on first run)
  3. Fetch filings — pulls each company’s submissions from data.sec.gov, filters to 10-K/10-Q/8-K filed since since, deduplicates against seen accessions. Fetches ~500-char narrative excerpts for up to 3 new filings per company
  4. Fetch news — Yahoo Finance RSS per ticker with 1s throttle, deduplicates by sha256(url)[:16]
  5. Build bundle — ranks filings + news by recency + type_weight (10-K=0.5, 10-Q=0.3, 8-K=0.15, news=0.1; recency = 1/(1+days_old)), applies a quota split (~55% filings, rest to news), caps at 18 items
  6. Update state — marks all new items seen, bumps last_run, saves
  7. Write outputs — versioned runs/<run_id>/<sector>.bundle.json, run.log, and run.summary.json

Phase 2 — Agent Layer (Researcher + Critic)

Section titled “Phase 2 — Agent Layer (Researcher + Critic)”

Two single-shot LLM agents wrapped in a deterministic revision-loop controller:

  • Researcher agent — reads the evidence bundle and produces structured claims, each citing a source_id from the bundle. Uses forced structured output via Anthropic’s tool-call API (tool_choice forced, strict: True) — the model emits a tool_use block guaranteed to match the Pydantic schema, never free text.
  • Critic agent — receives the bundle and the Researcher’s draft, checks every claim against the evidence, and returns a verdict: approve or revise with structured issues (each carrying claim_id, issue_type, explanation).
  • Revision-loop controller — deterministic state machine that decides when to cycle. Caps at 2 revision cycles (max 3 Researcher + 3 Critic calls). On cap exhaustion, the brief is marked flagged — never silently shipped as approved.
  1. Server-side — Anthropic’s API guarantees the model’s tool_use.input matches the input_schema (strict mode + forced tool choice)
  2. Client-side — Pydantic model_validate catches enum violations, empty cited source IDs, and verdict/issue consistency. The controller additionally verifies every cited_source_id exists in the bundle’s source_id set
flowchart TD
    S([load evidence bundle]) --> C0{cycle 0}
    C0 --> R0[Researcher — bundle only]
    R0 --> CR0[Critic — bundle + draft]
    CR0 --> V0{verdict?}
    V0 -- approve --> DONE0([approved, 0 revisions])
    V0 -- revise --> C1{cycle 1}
    C1 --> R1[Researcher — bundle + draft + critic issues]
    R1 --> CR1[Critic — bundle + revised draft]
    CR1 --> V1{verdict?}
    V1 -- approve --> DONE1([approved, 1 revision])
    V1 -- revise --> C2{cycle 2 = max}
    C2 --> R2[Researcher — final revision attempt]
    R2 --> CR2[Critic — final review]
    CR2 --> V2{verdict?}
    V2 -- approve --> DONE2([approved, 2 revisions])
    V2 -- revise --> FLAG([flagged — revision cap exhausted, last draft kept])
    R0 -. schema fail .-> FAIL([flagged — schema validation failure])

Key properties:

  • Deterministic — the controller decides when to cycle, not the agents
  • Stateless revisions — each cycle is a fresh messages.create call. Bundle, previous draft, and Critic output are re-supplied via the revision-turn prompt template. Nothing persists in context.
  • The Critic never flags — overall_verdict is exactly approve | revise. Flagging is assigned by the controller only, when the cap is hit or schema validation fails.
  • Never silently ships — on cap exhaustion, the brief is marked flagged with flag_reason="revision_cap_exhausted" and the last draft is kept

flowchart LR
    subgraph DATA["Data Pipeline (Deterministic, No LLM)"]
        EDGAR[SEC EDGAR<br/>10-K/10-Q/8-K filings]
        NEWS[Yahoo Finance RSS<br/>news per ticker]
        STATE[State Manager<br/>delta-only, per-sector]
        BUNDLE[Evidence Bundle<br/>ranked, quota-split JSON]
        EDGAR --> BUNDLE
        NEWS --> BUNDLE
        STATE --> BUNDLE
    end

    subgraph AGENTS["Agent Layer (LLM)"]
        RESEARCHER[Researcher Agent<br/>structured claims + citations]
        CRITIC[Critic Agent<br/>verdict: approve / revise]
        LOOP[Revision Loop Controller<br/>max 2 cycles, deterministic]
        BUNDLE --> RESEARCHER
        RESEARCHER --> CRITIC
        CRITIC --> LOOP
        LOOP -- revise --> RESEARCHER
    end

    subgraph OUTPUT["Output"]
        BRIEF[Audited Brief JSON<br/>full audit trail]
        LOOP --> BRIEF
    end
flowchart TB
    subgraph LOCAL["Local Machine"]
        APP[MarketBrief Application<br/>Python 3.12, uv-managed]
        CACHE[Company Cache<br/>weekly TTL]
        STATE2[State Files<br/>per-sector JSON]
        RUNS[Run Outputs<br/>versioned by UTC timestamp]
    end

    subgraph EXT["External APIs"]
        SEC[data.sec.gov<br/>SEC EDGAR]
        YAHOO[Yahoo Finance RSS]
        ANTHROPIC[Anthropic API<br/>Claude Sonnet 4.5]
    end

    APP -->|token-bucket limiter<br/>10 req/s| SEC
    APP -->|1s throttle per ticker| YAHOO
    APP -->|forced tool-call<br/>strict: True| ANTHROPIC
    APP --> CACHE
    APP --> STATE2
    APP --> RUNS

Layer Technology Details
Language Python 3.12 uv-managed dependencies
LLM Claude Sonnet 4.5 Forced structured output via tool-call API, temperature 0.0
Schema validation Pydantic ResearcherOutput, CriticOutput — single source of truth for both API schema and post-parse validation
Data sources SEC EDGAR + Yahoo Finance RSS 10-K/10-Q/8-K filings + per-ticker news feeds
Rate limiting Token-bucket (custom) SEC-compliant User-Agent enforced at construction, 10 req/s, exponential backoff on 429/5xx
State management Per-sector JSON Delta-only — re-running yields only items new since the prior run
Config YAML (config/pipeline.yaml) Sectors, quotas, rate limits, agent settings
Testing pytest 29 unit tests (no network) + integration tests (real API)
  • Software & Services — SIC 7372/7373/7374/7379, priority tickers: MSFT, ORCL, ADBE, CRM, NOW, INTU, IBM
  • Semiconductors & Equipment — SIC 3674, priority tickers: NVDA, AMD, INTC, AVGO, TXN, QCOM, MU, AMAT, LRCX, KLAC

Both sectors are configurable — additional sectors can be added via config/pipeline.yaml.


Every draft, every Critic verdict with full issue detail, the revision count, and the final approved-or-flagged status — all persisted for review.


  • Deterministic data pipeline, then LLM — the data pipeline has zero LLM code. Evidence bundles are the handoff artifact. This separation means the data layer can be tested, debugged and re-run independently of API costs.
  • Forced structured output, not free-text parsing — the Researcher and Critic use Anthropic’s tool-call API with strict: True and forced tool_choice. The model emits a tool_use block guaranteed to match the Pydantic schema — no regex parsing, no prompt-engineering for format compliance.
  • Two-layer schema enforcement — server-side guarantee from Anthropic + client-side Pydantic validation. The controller additionally verifies every cited id exists in the bundle.
  • The Critic never flags — is exactly approve | revise. Flagging is assigned by the controller only, when the revision cap is hit or schema validation fails. This keeps the agent’s role clean and the control flow deterministic.
  • Never silently ships a bad brief — on cap exhaustion, the brief is marked flagged, the reason is recorded, and the last draft is kept for human review. No silent approval.
  • Delta-only state — re-running the pipeline yields only items new since the prior run. State is the source of truth, not a re-scan of everything.

This project explores how to build AI systems where trust is engineered into the architecture, not bolted on. The Researcher-Critic pattern with a deterministic revision loop is a design I keep returning to — it separates the generation and validation concerns cleanly, makes the control flow predictable, and produces an audit trail that a human can actually review.

The deeper lesson is about the boundary between deterministic code and LLM calls: the data pipeline is deterministic and testable, the agent calls are structured and validated, and the controller is a state machine. Each layer has a different failure mode and a different way to verify it.