Your finance RAG keeps nailing demos – until real filings show up
You’ve built a retrieval system over 10-Ks and earnings. It answers “What was revenue last year?” cleanly in the lab. Then a real analyst asks for FY2022 fixed-asset turnover or whether operating margins stayed consistent after one-offs. Suddenly numbers drift, units flip, or the model politely refuses.
That gap is expensive. Lab demos don’t lose money; production answers can.
FinanceBench is the open-book financial QA benchmark built for that exact failure mode. Real U.S. public filings. Free 150-case sample. One evening of scoring tells you whether the stack is shippable.
What FinanceBench actually tests
The November 2023 paper (arXiv:2311.11944) from Patronus AI (Islam et al.) defines the full suite: 10,231 question-answer-evidence triplets, 40 U.S.-listed companies, 361 filings (mostly 10-Ks, plus 10-Qs, 8-Ks, earnings) spanning 2015-2023.
Questions stay clear on purpose – a minimum bar, not a torture test. Three flavors show up: metrics-generated (pull or compute a number), domain-relevant (25 generic analyst prompts – liquidity, capital intensity, dividend trends – applied per company), and novel-generated (company-specific but realistic). About two-thirds need numerical reasoning. The rest mix extraction and light logic.
Gold answer, optional justification, evidence with page numbers and full-page text – each row ships complete. Page numbers are zero-indexed; that detail bites people on day one. The 150 human-scored cases live on GitHub and Hugging Face under cc-by-nc-4.0. PDFs sit in /pdfs. Want the full 10k? Email [email protected] – public leaderboards for every new model still do not exist.
Practical setup: load the sample and score something in one sitting
Clone the repo or pull the HF set. Two JSONL files merge on doc_name:
import pandas as pd
df_q = pd.read_json("data/financebench_open_source.jsonl", lines=True)
df_m = pd.read_json("data/financebench_document_information.jsonl", lines=True)
df = pd.merge(df_q, df_m, on="doc_name")
print(df[["company", "question", "answer", "question_type"]].head())
Rows look like 3M FY2018 capex from the cash-flow statement ($1,577 million) or Activision fixed-asset turnover that forces a multi-year PP&E average. Match PDF via doc_name (example: BOEING_2022_10K).
Minimal eval loop:
- Grab 20-30 rows spanning metrics-generated and domain-relevant.
- Pair each question with either retrieved chunks from your store over that PDF, or gold
evidence_text/ full page for an oracle baseline. - Generate with your model.
- Score hard: exact or fuzzy match to gold (Patronus cookbook uses fuzzy). Bucket correct / incorrect / refused – same three labels the paper used.
Repo extras: evaluation_playground.ipynb and pre-computed outputs under /results.
Pro tip: Log the
evidence_page_numyou actually retrieved. Miss the gold page and the failure is retrieval, not the LLM. Zero-index off-by-one is the classic first-hour bug.
Map the paper configs to your real pipeline
Human review on n=150 (2,400 answers) still anchors baselines – as of the 2023 paper numbers:
| Configuration | Approx. correct (GPT-4-Turbo) | What it means for you |
|---|---|---|
| Closed-book | ~9% | Model memory alone is useless on fresh filings |
| Shared vector store (all docs) | ~19% | Classic enterprise default – cross-doc noise wrecks recall |
| Per-filing vector store | ~50% | Isolation helps; still half wrong or refused |
| Long-context (whole relevant filing) | ~79% | Stronger, slow and costly on 100+ page 10-Ks |
| Oracle (gold pages only) | ~85% | Ceiling with perfect retrieval; leftover mistakes are arithmetic |
Shared-store style setups – the ones most demos ship – landed GPT-4-Turbo around 81% incorrect or refused. Llama-2 variants hallucinated more; GPT-4 refused more. Neither pattern is safe unsupervised.
One giant index across dozens of tickers? Expect the ~19% regime. Split by ticker or filing year and re-measure. That single change is the biggest win most teams skip.
Why does perfect page context still leave a ~15% hole? Units, signs, multi-step ratios. Retrieval is only half the story – and the boring half once isolation works.
Advanced: push past the 50% wall
Isolate retrieval first. Then hunt numerical failures: millions vs billions, sign flips, CAGR / turnover / margin-after-one-offs. Attach a calculator tool or force chain-of-thought with explicit intermediate lines. Re-score the same 30 cases. Don’t change five variables at once.
Tree-style / outline-reasoning indexes have posted much higher accuracy on the public sample under single-database protocols (claims as of their 2025-2026 writeups). Re-run the identical 150 yourself. Protocol drift is real; the 2023 vector baselines remain the reference for classic RAG failure modes.
Metadata filters (company, year, doc_type) inside the query beat post-filtering. Community reports show large recall lifts when the filter is part of search, not an afterthought.
Honest limits you should budget for
n=150 is small. A few points between models is noise. The full 10,231 set exists; public exact scores for every modern model do not. Full reproduction still goes through Patronus.
Single-turn, single-company only. Real analyst work is multi-hop, multi-year, comparative. Some items are simplistic by design – the paper calls this a minimum standard. HF sample license is non-commercial. Long-context wins carry the latency warning the authors already flagged as unrealistic for big enterprise docs.
If your eval only checks “did it retrieve something,” you will over-estimate production safety.
FAQ
Is the full FinanceBench free?
No. Only the 150-case sample (PDFs + some results) is public on GitHub/HF. Contact Patronus for the 10,231-question set or platform access.
Can I train on it?
Evaluation suite, not a training corpus. Open sample on Hugging Face is cc-by-nc-4.0. Keep it held-out.
My RAG scores 70%+ on the sample – am I done?
Check the protocol before celebrating. Per-doc isolation or shared index? Refusals and numerical exactness scored the paper’s way? Gold page actually in top-k? Run the same questions with a calculator on and with deliberately noisy retrieval. Classic setups top out near the oracle ~85%. Near-perfect claims need a published, re-runnable recipe.
Clone the repo tonight. Load the JSONL. Pick ten metrics-generated questions. Score your pipeline against gold. That one run beats another week of demo polish.