Skip to content

Chapter 9 of 10 · 29 entries

9. Evaluation Benchmarks and Arenas ​

Measure research QA, trading decisions, safety, and live arena performance. Without evals, finance agents are demos.

Research and open-book QA ​

  • PIXIU FinMA + FinBen financial LLM evaluation suite.
  • FinBen Holistic financial LLM benchmark from The-FinAI.
  • financebench Open-book financial QA benchmark over real filings.
  • finance-agent Tool-using finance agent research benchmark (arXiv:2508.00828).
  • finance-agent-v2 Finance Agent benchmark v2 with EDGAR, web, and price tools.
  • FinSearchComp Expert financial search and reasoning benchmark (ByteDance Seed lineage).
  • BizFinBench Business-driven real-world financial LLM benchmark.
  • FinEval Chinese financial LLM evaluation suite. CN
  • MME-Finance Multimodal finance benchmark for expert-level understanding.
  • FinToolBench 760 executable financial tools and 295 queries for evaluating tool-using agents.
  • FinMTM Multi-turn multimodal benchmark for financial reasoning and agents.
  • finchain Symbolic benchmark for verifiable chain-of-thought financial reasoning.
  • big-finance-benchmark Reference harness for BigFinanceBench, a workflow-grounded financial research benchmark.
  • CNFinBench OpenCompass benchmark for high-stakes financial agent tasks. CN
  • FinanceBenchmark Microsoft benchmark for finance agents in M365 Copilot (~300 tasks with ERP and web tools).

Decision and trading benchmarks ​

  • INVESTOR-BENCH InvestorBench: LLM agent financial decision benchmark (ACL 2025).
  • QuantitativeFinance-Bench State-aware interactive quant agent benchmark on Harbor.
  • stockbench Controlled benchmark asking whether LLM agents can trade profitably.
  • DeepFund NeurIPS'25 pilot for fund-investment style agent evaluation.
  • QuantCode-Bench Benchmark for executable algorithmic trading code generation.
  • TwinMarket LLM-driven market simulation for agent interaction studies.

Live arenas and public track records ​

  • live-trade-bench Live evaluation of trading agents.
  • Agent_Market_Arena Live multi-market trading arena for LLM agents.
  • alpha-arena Alpha arena competition harness for trading agents.
  • LLM-Trading-Lab Public experiment where ChatGPT manages a small real-money micro-cap portfolio, with logs and code. live
  • nof0 Open-source AI trading arena in the style of nof1 Alpha Arena. CN
  • continuous-record-llm-trading-agents Data and artifacts from a continuous-record study of what LLM trading agents actually do. paper

Safety and trustworthiness ​

Pitfall Tip ​

Leaderboard PnL without costs, latency, or leakage controls is not evidence. Prefer benchmarks that publish contamination policy and tool access rules.

For research and education only. Not investment advice.