# 9. 评测基准与竞技场

> 第 9 章，共 10 章 · 29 条 · 网页版: https://finance-agents.dev/zh/chapters/9-evaluation-benchmarks-and-arenas

衡量研究问答、交易决策、安全性和实盘竞技场表现。没有评测，金融 agent 就只是演示。

## 研究与开卷问答

- [PIXIU](https://github.com/The-FinAI/PIXIU) - FinMA + FinBen 金融 LLM 评测套件。
- [FinBen](https://github.com/The-FinAI/FinBen) - The-FinAI 的综合性金融 LLM 基准。
- [financebench](https://github.com/patronus-ai/financebench) - 基于真实财报的开卷金融问答基准。
- [finance-agent](https://github.com/vals-ai/finance-agent) - 带工具调用的金融研究 agent 基准（arXiv:2508.00828）。
- [finance-agent-v2](https://github.com/vals-ai/finance-agent-v2) - Finance Agent 基准 v2，含 EDGAR、网页与价格工具。
- [FinSearchComp](https://github.com/randomtutu/FinSearchComp) - 专家级金融搜索与推理基准（字节 Seed 相关）。
- [BizFinBench](https://github.com/HiThink-Research/BizFinBench) - 面向真实业务场景的金融 LLM 基准。
- [FinEval](https://github.com/SUFE-AIFLM-Lab/FinEval) - 🇨🇳 中文金融 LLM 评测套件。
- [MME-Finance](https://github.com/HiThink-Research/MME-Finance) - 面向专家级理解的多模态金融基准。
- [FinToolBench](https://github.com/Double-wk/FinToolBench) - 760 个可执行金融工具、295 个查询，评测工具调用 agent。
- [FinMTM](https://github.com/HiThink-Research/FinMTM) - 多轮多模态金融推理与 agent 基准。
- [finchain](https://github.com/mbzuai-nlp/finchain) - 可验证思维链金融推理的符号化基准。
- [big-finance-benchmark](https://github.com/Rogo-Technologies/big-finance-benchmark) - BigFinanceBench 参考评测框架，基于真实工作流的金融研究基准。
- [CNFinBench](https://github.com/open-compass/CNFinBench) - 🇨🇳 OpenCompass 出品的高风险金融 agent 任务基准。
- [FinanceBenchmark](https://github.com/microsoft/FinanceBenchmark) - 微软 M365 Copilot 财务 agent 基准（约 300 个任务，含 ERP 与网页工具）。

## 决策与交易基准

- [INVESTOR-BENCH](https://github.com/felis33/INVESTOR-BENCH) - InvestorBench：LLM agent 金融决策基准（ACL 2025）。
- [QuantitativeFinance-Bench](https://github.com/QF-Bench/QuantitativeFinance-Bench) - 在 Harbor 上的状态感知交互式量化 agent 基准。
- [stockbench](https://github.com/ChenYXxxx/stockbench) - 受控环境下检验 LLM agent 能否交易盈利的基准。
- [DeepFund](https://github.com/HKUSTDial/DeepFund) - NeurIPS'25 基金投资风格的 agent 评测试点。
- [QuantCode-Bench](https://github.com/LimexAILab/QuantCode-Bench) - 可执行算法交易代码生成基准。
- [TwinMarket](https://github.com/FreedomIntelligence/TwinMarket) - 用于研究 agent 交互的 LLM 驱动市场仿真。

## 实盘竞技场与公开战绩

- [live-trade-bench](https://github.com/ulab-uiuc/live-trade-bench) - 交易 agent 的实盘/在线评测。
- [Agent_Market_Arena](https://github.com/The-FinAI/Agent_Market_Arena) - 面向 LLM agent 的多市场实盘竞技场。
- [alpha-arena](https://github.com/AmadeusGB/alpha-arena) - 交易 agent 的 Alpha 竞技场竞赛 harness。
- [LLM-Trading-Lab](https://github.com/LuckyOne7777/LLM-Trading-Lab) - 公开实验：ChatGPT 管理一笔小额真实资金的微盘股组合，附日志与代码。(live)
- [nof0](https://github.com/wquguru/nof0) - 🇨🇳 开源 AI 交易竞技场，仿 nof1 Alpha Arena。
- [continuous-record-llm-trading-agents](https://github.com/ProjectDXAI/continuous-record-llm-trading-agents) - 对 LLM 交易 agent 实际行为做连续记录研究的数据与产物。(paper)

## 安全与可信

- [awesome-financial-llm-trustworthiness-benchmarks](https://github.com/FDU-INS/awesome-financial-llm-trustworthiness-benchmarks) - 金融 LLM 可信度相关基准的精选清单。
- [TradeTrap](https://github.com/Yanlewen/TradeTrap) - 检验 LLM 交易 agent 在扰动下是否可靠、忠实。(paper)

## 踩坑提示

没有成本、延迟和防泄漏控制的排行榜收益不能当证据。优先选择公开污染策略和工具访问规则的基准。

---

不构成投资建议；收录不代表推荐。 CC0 1.0 · https://github.com/aowang-ai/awesome-finance-agents

由 Ao Wang（王奥，https://aowang.ai）维护。
