Chapter 9 of 10 · 29 entries
9. Evaluation Benchmarks and Arenas
Measure research QA, trading decisions, safety, and live arena performance. Without evals, finance agents are demos.
Research and open-book QA
- PIXIU FinMA + FinBen financial LLM evaluation suite.
- FinBen Holistic financial LLM benchmark from The-FinAI.
- financebench Open-book financial QA benchmark over real filings.
- finance-agent Tool-using finance agent research benchmark (arXiv:2508.00828).
- finance-agent-v2 Finance Agent benchmark v2 with EDGAR, web, and price tools.
- FinSearchComp Expert financial search and reasoning benchmark (ByteDance Seed lineage).
- BizFinBench Business-driven real-world financial LLM benchmark.
- FinEval Chinese financial LLM evaluation suite. CN
- MME-Finance Multimodal finance benchmark for expert-level understanding.
- FinToolBench 760 executable financial tools and 295 queries for evaluating tool-using agents.
- FinMTM Multi-turn multimodal benchmark for financial reasoning and agents.
- finchain Symbolic benchmark for verifiable chain-of-thought financial reasoning.
- big-finance-benchmark Reference harness for BigFinanceBench, a workflow-grounded financial research benchmark.
- CNFinBench OpenCompass benchmark for high-stakes financial agent tasks. CN
- FinanceBenchmark Microsoft benchmark for finance agents in M365 Copilot (~300 tasks with ERP and web tools).
Decision and trading benchmarks
- INVESTOR-BENCH InvestorBench: LLM agent financial decision benchmark (ACL 2025).
- QuantitativeFinance-Bench State-aware interactive quant agent benchmark on Harbor.
- stockbench Controlled benchmark asking whether LLM agents can trade profitably.
- DeepFund NeurIPS'25 pilot for fund-investment style agent evaluation.
- QuantCode-Bench Benchmark for executable algorithmic trading code generation.
- TwinMarket LLM-driven market simulation for agent interaction studies.
Live arenas and public track records
- live-trade-bench Live evaluation of trading agents.
- Agent_Market_Arena Live multi-market trading arena for LLM agents.
- alpha-arena Alpha arena competition harness for trading agents.
- LLM-Trading-Lab Public experiment where ChatGPT manages a small real-money micro-cap portfolio, with logs and code. live
- nof0 Open-source AI trading arena in the style of nof1 Alpha Arena. CN
- continuous-record-llm-trading-agents Data and artifacts from a continuous-record study of what LLM trading agents actually do. paper
Safety and trustworthiness
- awesome-financial-llm-trustworthiness-benchmarks Curated benchmarks for financial LLM trustworthiness.
- TradeTrap Tests whether LLM trading agents stay reliable and faithful under perturbation. paper
Pitfall Tip
Leaderboard PnL without costs, latency, or leakage controls is not evidence. Prefer benchmarks that publish contamination policy and tool access rules.