Registry section
Datasets
Data for training and evaluating trading agents.Filter and compare in Browse →

Binance's official archive of free, unauthenticated historical market data — daily and monthly ZIPs of klines, trades and aggregate trades for spot and futures across 3,695 spot symbols.
openklines from 1s to 1mo, plus trades and aggregate tradesThree top-level trees (spot, futures, option), served as per-symbol daily and monthly ZIPscryptoadded 2026-08-17 · unknown · external- ECTSumDataset
A benchmark of 2,425 long earnings call transcripts paired with expert-written telegram-style bullet-point summaries drawn from Reuters articles.
open2,425 document–summary pairsequitiesadded 2026-08-15 · GPL-3.0 · external - EDGAR-CORPUSDataset
A corpus of 220,375 SEC 10-K annual reports from 1993 to 2020, split into their individual item sections and released as JSON.
openannual~40.7 GB (176,289 train / 22,050 validation / 22,036 test in the full config)equitiesadded 2026-08-15 · Apache-2.0 · external - Fama/French Data LibraryDataset
Kenneth French's library of the Fama/French factor return series — 3-factor, 5-factor, momentum and reversal, plus hundreds of sorted portfolios — as CSV archives reaching back to July 1926.
opendaily, weekly, monthly and annualSmall — individual factor archives are 5-13 KB zippedequitiesadded 2026-08-17 · unknown · external The first public benchmark dataset of high-frequency limit order book data: about 4,000,000 samples from five NASDAQ Nordic stocks over ten days, with 144 features and mid-price labels.
openevent-based limit order book messages~4,000,000 time-series samples in a single 1.86 GB BenchmarkDatasets.zipequitiestrading: advancedadded 2026-08-17 · CC-BY-4.0 · externalA human-annotated sentiment dataset of 4,840 English sentences from financial news, labelled positive, negative, or neutral from an investor's perspective.
open4,840 sentences in four agreement-level configurationsequitiesadded 2026-08-15 · CC-BY-NC-SA-3.0 · external- FNSPIDDataset
A time-series dataset pairing 29.7 million stock price records with 15.7 million financial news articles for 4,775 companies in the S&P 500 universe, 1999–2023.
opendaily prices with timestamped news29.7M price records + 15.7M news records (~30 GB)equitiesadded 2026-08-15 · CC-BY-NC-4.0 · external 2,480 hand-annotated sentences from FOMC speeches, meeting minutes and press-conference transcripts labelled hawkish, dovish or neutral, released with the ACL 2023 Trillion Dollar Words paper.
open2,480 labelled sentences (~526 KB)ratesequitiesadded 2026-08-17 · CC-BY-NC-4.0 · external- Hyperliquid L4 Order BookDataset
A 195.3 GB Zenodo release of level-4 order flow from the Hyperliquid perpetual futures DEX — including rejected orders for BTC, ETH and SOL, book diffs, and trades across 250+ contracts.
openevent-based, nanosecond timestamps195.3 GB across 15 files, largest single archive 49.6 GBcryptotrading: advancedprogramming: advancedsetup: advancedadded 2026-08-17 · CC-BY-4.0 · external - JKP Global Factor DataDataset
Jensen, Kelly and Pedersen's global factor dataset: 153 firm characteristics across 13 themes for 93 countries and 4 regions, published as downloadable factor and portfolio return series.
openmonthly and dailyIndividual factor zips are tens of kilobytes; the gated CTF dataset is roughly 1.1 GB of Parquetequitiesadded 2026-08-17 · CC-BY-NC-4.0 (data); MIT (pipeline code) · external - LOBSTERDataset
Nanosecond-timestamped limit order book reconstructions from NASDAQ Historical TotalView-ITCH for the entire NASDAQ universe since 27 June 2007, sold by subscription.
commercialevent-based (per-order messages, millisecond to nanosecond precision)One message CSV and one orderbook CSV per ticker per trading day (~74 MB uncompressed for one stock, one day, 10 levels)equitiestrading: advancedadded 2026-08-17 · proprietary · external Per-filing summary records for every SEC 10-K/10-Q variant filed on EDGAR since 1993, with Loughran-McDonald sentiment word counts precomputed and keyed by CIK and filing date.
openevent-based (one row per filing)1,249,506 rows x 26 columns (~188 MB)equitiesadded 2026-08-17 · non-commercial (SRAF terms) · external- Open Source Asset PricingDataset
An open reproduction of the cross-sectional return-predictability literature: 212 predictor signals plus 114 placebos, released as portfolio returns, firm characteristics and construction code.
openmonthly and daily portfolio returns331 signal rows in SignalDoc.csv; ~1.6 GB zipped CSV for firm-level characteristicsequitiesadded 2026-08-17 · GPL-2.0 · external - Polymarket-v1Dataset
The complete on-chain trade tape and contract lifecycle for Polymarket v1 on Polygon — 1.20 billion maker-taker fills plus cleaned analysis layers and CTF lifecycle events, in Parquet.
openevent-based, one row per matched fill1.20B fills in the raw tape; ~52.7 GB of Parquet across four configsprediction-marketsprogramming: advancedadded 2026-08-17 · CC-BY-4.0 · external The SEC's quarterly release of numeric data from the primary financial statements of every XBRL filing made to the Commission, January 2009 through March 2026.
openquarterly69 quarterly ZIPs, from 13.22 KB (2009 Q1) to 121.92 MB (2025 Q3)equitiesadded 2026-08-17 · US Government work · external- StockNet datasetDataset
Two years of tweets and Yahoo Finance prices for 88 high-cap stocks across 9 sectors, released with Xu and Cohen's ACL 2018 paper and the source of the reused ACL18 stock-movement task.
opendaily55,141 files, ~460 MBequitiesadded 2026-08-17 · MIT · external - TeraflopAI SEC-EDGARDataset
A corpus of 8,055,455 SEC EDGAR filings — 10-K, 10-Q, 8-K, S-1, S-8, 20-F, 144 and Forms 3/4/5 — carrying raw filing bytes, parsed plaintext and filer, accession and period metadata.
openevent-based, one row per filing8,055,455 rows, ~295 GB of Parquet on Hugging Face (43.7B tokens)equitiesadded 2026-08-17 · Apache-2.0 · external
Missing something? Add it — no YAML required.