Financial Market Brief Generator
An AI system that generates audited daily market briefs from SEC filings and news — with a Researcher agent, a Critic agent, and a deterministic revision loop
Problem
Section titled “Problem”Tracking specific sectors of the financial markets is a daily grind — scanning SEC filings, reading earnings reports, monitoring news feeds, then synthesizing everything into a coherent summary. Most existing tools give you everything; what you need is a focused brief on the sectors you actually care about, with evidence you can trace back to the source.
The harder problem: how do you trust that an AI-generated brief is grounded in real data and not hallucinated? Generating text is easy — ensuring every claim is backed by a cited source, reviewed by a second agent, and revised when it fails review is the actual engineering challenge.
Approach
Section titled “Approach”The system is built in two phases — a deterministic data pipeline that produces evidence bundles, and an agent layer (Researcher + Critic) that turns those bundles into audited briefs through a revision loop.
Phase 1 — Data Pipeline (No LLM)
Section titled “Phase 1 — Data Pipeline (No LLM)”A deterministic Python pipeline that turns a GICS sector into a versioned evidence bundle (JSON) of new SEC filings and news items since the last run. No LLM code — pure data plumbing.
7-step pipeline per sector:
- Resolve companies — weekly-cached company list per sector (priority tickers first, then EDGAR full-text search by SIC code)
- Load state — per-sector JSON tracking last run timestamp, seen SEC accession numbers, and seen news URL hashes. Computes the
sincecutoff (last run, ornow − 7 dayson first run) - Fetch filings — pulls each company’s submissions from
data.sec.gov, filters to 10-K/10-Q/8-K filed sincesince, deduplicates against seen accessions. Fetches ~500-char narrative excerpts for up to 3 new filings per company - Fetch news — Yahoo Finance RSS per ticker with 1s throttle, deduplicates by
sha256(url)[:16] - Build bundle — ranks filings + news by
recency + type_weight(10-K=0.5, 10-Q=0.3, 8-K=0.15, news=0.1; recency =1/(1+days_old)), applies a quota split (~55% filings, rest to news), caps at 18 items - Update state — marks all new items seen, bumps
last_run, saves - Write outputs — versioned
runs/<run_id>/<sector>.bundle.json,run.log, andrun.summary.json
Phase 2 — Agent Layer (Researcher + Critic)
Section titled “Phase 2 — Agent Layer (Researcher + Critic)”Two single-shot LLM agents wrapped in a deterministic revision-loop controller:
- Researcher agent — reads the evidence bundle and produces structured claims, each citing a
source_idfrom the bundle. Uses forced structured output via Anthropic’s tool-call API (tool_choiceforced,strict: True) — the model emits atool_useblock guaranteed to match the Pydantic schema, never free text. - Critic agent — receives the bundle and the Researcher’s draft, checks every claim against the evidence, and returns a verdict:
approveorrevisewith structured issues (each carryingclaim_id,issue_type,explanation). - Revision-loop controller — deterministic state machine that decides when to cycle. Caps at 2 revision cycles (max 3 Researcher + 3 Critic calls). On cap exhaustion, the brief is marked
flagged— never silently shipped as approved.
Schema enforcement — two layers
Section titled “Schema enforcement — two layers”- Server-side — Anthropic’s API guarantees the model’s
tool_use.inputmatches theinput_schema(strict mode + forced tool choice) - Client-side — Pydantic
model_validatecatches enum violations, empty cited source IDs, and verdict/issue consistency. The controller additionally verifies everycited_source_idexists in the bundle’ssource_idset
Revision-loop flow
Section titled “Revision-loop flow”flowchart TD
S([load evidence bundle]) --> C0{cycle 0}
C0 --> R0[Researcher — bundle only]
R0 --> CR0[Critic — bundle + draft]
CR0 --> V0{verdict?}
V0 -- approve --> DONE0([approved, 0 revisions])
V0 -- revise --> C1{cycle 1}
C1 --> R1[Researcher — bundle + draft + critic issues]
R1 --> CR1[Critic — bundle + revised draft]
CR1 --> V1{verdict?}
V1 -- approve --> DONE1([approved, 1 revision])
V1 -- revise --> C2{cycle 2 = max}
C2 --> R2[Researcher — final revision attempt]
R2 --> CR2[Critic — final review]
CR2 --> V2{verdict?}
V2 -- approve --> DONE2([approved, 2 revisions])
V2 -- revise --> FLAG([flagged — revision cap exhausted, last draft kept])
R0 -. schema fail .-> FAIL([flagged — schema validation failure])
Key properties:
- Deterministic — the controller decides when to cycle, not the agents
- Stateless revisions — each cycle is a fresh
messages.createcall. Bundle, previous draft, and Critic output are re-supplied via the revision-turn prompt template. Nothing persists in context. - The Critic never flags —
overall_verdictis exactlyapprove | revise. Flagging is assigned by the controller only, when the cap is hit or schema validation fails. - Never silently ships — on cap exhaustion, the brief is marked
flaggedwithflag_reason="revision_cap_exhausted"and the last draft is kept
Architecture
Section titled “Architecture”System architecture
Section titled “System architecture”flowchart LR
subgraph DATA["Data Pipeline (Deterministic, No LLM)"]
EDGAR[SEC EDGAR<br/>10-K/10-Q/8-K filings]
NEWS[Yahoo Finance RSS<br/>news per ticker]
STATE[State Manager<br/>delta-only, per-sector]
BUNDLE[Evidence Bundle<br/>ranked, quota-split JSON]
EDGAR --> BUNDLE
NEWS --> BUNDLE
STATE --> BUNDLE
end
subgraph AGENTS["Agent Layer (LLM)"]
RESEARCHER[Researcher Agent<br/>structured claims + citations]
CRITIC[Critic Agent<br/>verdict: approve / revise]
LOOP[Revision Loop Controller<br/>max 2 cycles, deterministic]
BUNDLE --> RESEARCHER
RESEARCHER --> CRITIC
CRITIC --> LOOP
LOOP -- revise --> RESEARCHER
end
subgraph OUTPUT["Output"]
BRIEF[Audited Brief JSON<br/>full audit trail]
LOOP --> BRIEF
end
Pipeline runtime topology
Section titled “Pipeline runtime topology”flowchart TB
subgraph LOCAL["Local Machine"]
APP[MarketBrief Application<br/>Python 3.12, uv-managed]
CACHE[Company Cache<br/>weekly TTL]
STATE2[State Files<br/>per-sector JSON]
RUNS[Run Outputs<br/>versioned by UTC timestamp]
end
subgraph EXT["External APIs"]
SEC[data.sec.gov<br/>SEC EDGAR]
YAHOO[Yahoo Finance RSS]
ANTHROPIC[Anthropic API<br/>Claude Sonnet 4.5]
end
APP -->|token-bucket limiter<br/>10 req/s| SEC
APP -->|1s throttle per ticker| YAHOO
APP -->|forced tool-call<br/>strict: True| ANTHROPIC
APP --> CACHE
APP --> STATE2
APP --> RUNS
Tech stack
Section titled “Tech stack”| Layer | Technology | Details |
|---|---|---|
| Language | Python 3.12 | uv-managed dependencies |
| LLM | Claude Sonnet 4.5 | Forced structured output via tool-call API, temperature 0.0 |
| Schema validation | Pydantic | ResearcherOutput, CriticOutput — single source of truth for both API schema and post-parse validation |
| Data sources | SEC EDGAR + Yahoo Finance RSS | 10-K/10-Q/8-K filings + per-ticker news feeds |
| Rate limiting | Token-bucket (custom) | SEC-compliant User-Agent enforced at construction, 10 req/s, exponential backoff on 429/5xx |
| State management | Per-sector JSON | Delta-only — re-running yields only items new since the prior run |
| Config | YAML (config/pipeline.yaml) |
Sectors, quotas, rate limits, agent settings |
| Testing | pytest | 29 unit tests (no network) + integration tests (real API) |
Configured sectors
Section titled “Configured sectors”- Software & Services — SIC 7372/7373/7374/7379, priority tickers: MSFT, ORCL, ADBE, CRM, NOW, INTU, IBM
- Semiconductors & Equipment — SIC 3674, priority tickers: NVDA, AMD, INTC, AVGO, TXN, QCOM, MU, AMAT, LRCX, KLAC
Both sectors are configurable — additional sectors can be added via config/pipeline.yaml.
Audit trail
Section titled “Audit trail”Every draft, every Critic verdict with full issue detail, the revision count, and the final approved-or-flagged status — all persisted for review.
Key design decisions
Section titled “Key design decisions”- Deterministic data pipeline, then LLM — the data pipeline has zero LLM code. Evidence bundles are the handoff artifact. This separation means the data layer can be tested, debugged and re-run independently of API costs.
- Forced structured output, not free-text parsing — the Researcher and Critic use Anthropic’s tool-call API with
strict: Trueand forcedtool_choice. The model emits atool_useblock guaranteed to match the Pydantic schema — no regex parsing, no prompt-engineering for format compliance. - Two-layer schema enforcement — server-side guarantee from Anthropic + client-side Pydantic validation. The controller additionally verifies every cited
idexists in the bundle. - The Critic never flags — is exactly
approve | revise. Flagging is assigned by the controller only, when the revision cap is hit or schema validation fails. This keeps the agent’s role clean and the control flow deterministic. - Never silently ships a bad brief — on cap exhaustion, the brief is marked
flagged, the reason is recorded, and the last draft is kept for human review. No silent approval. - Delta-only state — re-running the pipeline yields only items new since the prior run. State is the source of truth, not a re-scan of everything.
What I am learning
Section titled “What I am learning”This project explores how to build AI systems where trust is engineered into the architecture, not bolted on. The Researcher-Critic pattern with a deterministic revision loop is a design I keep returning to — it separates the generation and validation concerns cleanly, makes the control flow predictable, and produces an audit trail that a human can actually review.
The deeper lesson is about the boundary between deterministic code and LLM calls: the data pipeline is deterministic and testable, the agent calls are structured and validated, and the controller is a state machine. Each layer has a different failure mode and a different way to verify it.