athenara:~$ registry bench benchmarks/investorbench

InvestorBench

A benchmark for evaluating LLM-based agents on financial decision-making tasks across single stocks, cryptocurrencies, and ETFs.

#benchmark #llm-agent #decision-making #evaluation

added 2026-08-15 · MIT · external

$ git clone https://github.com/felis33/INVESTOR-BENCH
collected 3 task environments
investorbench::stock-trading
investorbench::crypto-trading
investorbench::etf-trading

InvestorBench pairs a task-agnostic LLM agent framework with standardized datasets and simulated market environments so different backbone models can be compared on the same footing. The agent framework decomposes into Brain, Perception, Profile, Memory, and Action modules, with data drawn from open sources and third-party APIs including Yahoo Finance and SEC EDGAR.

The authors evaluate thirteen LLMs as backbones across market environments, scored with standard quantitative finance metrics. Published at ACL 2025 main conference (arXiv:2412.18174); the official implementation is released at felis33/INVESTOR-BENCH under MIT.

trading [●●●··] moderate   ai [●●●●·] advanced   programming [●●●··] moderate   setup [●●●●●] expert

authors Haohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji, Zhiyang Deng, Yueru He, Yuechen Jiang, Zining Zhu, K.P. Subbalakshmi, Jimin Huang, Lingfei Qian, Xueqing Peng, Jordan W. Suchow, Qianqian Xie
origin external
license MIT
markets multi-asset

athenara:~$