Registry section
Benchmarks
Standardized evaluations for trading agents and financial models.Filter and compare in Browse →

- DeepFundBenchmark
A live, leakage-free benchmark that runs a multi-agent LLM fund-management workflow against market data postdating the models' training cutoffs.
stock-tradingportfolio-managementdecision-makingai: advancedsetup: advancedadded 2026-08-17 · MIT · external - FinBenBenchmark
A holistic open-source financial benchmark of 42 datasets spanning 24 tasks, including the first stock-trading evaluation for LLMs.
information-extractiontextual-analysisquestion-answeringtext-generationai: advancedadded 2026-08-15 · external - FINSABERBenchmark
A long-horizon backtesting benchmark that tests whether LLM investing agents, RL traders and classical strategies beat buy-and-hold on S&P 500 data from 2000 to 2025.
stock-tradingbacktestingai: advancedsetup: advancedadded 2026-08-17 · Apache-2.0 · external - FinSearchCompBenchmark
An open benchmark of 635 expert-written financial search-and-reasoning questions across Global and Greater China markets, with a released evaluation harness and public leaderboard.
time-sensitive-data-fetchingsimple-historical-lookupcomplex-historical-investigationadded 2026-08-17 · CC-BY-4.0 · external - InvestorBenchBenchmark
A benchmark for evaluating LLM-based agents on financial decision-making tasks across single stocks, cryptocurrencies, and ETFs.
stock-tradingcrypto-tradingetf-tradingai: advancedsetup: expertadded 2026-08-15 · MIT · external - LiveTradeBenchBenchmark
A live forward-testing harness that runs LLM trading agents against streaming market data across US equities and Polymarket prediction markets, using a portfolio-allocation abstraction.
stock-tradingprediction-market-tradingportfolio-allocationai: advancedadded 2026-08-17 · PolyForm-Noncommercial-1.0.0 · external - PIXIU / FLAREBenchmark
An open-source financial LLM resource bundling the FinMA instruction-tuned models, a 136K-sample instruction dataset, and the FLARE evaluation benchmark.
sentiment-analysisclassificationknowledge-extractionnumerical-reasoningai: advancedadded 2026-08-15 · MIT · external - StockBenchBenchmark
A frozen-window stock-trading benchmark in which LLM agents run a daily portfolio, analysis and trade loop over 20 DJIA constituents for 82 trading days of 2025.
stock-tradingdecision-makingai: advancedadded 2026-08-17 · Apache-2.0 · external
Missing something? Add it — no YAML required.