athenara:~$ registry open datasets/financial-phrasebank

Financial PhraseBank

A human-annotated sentiment dataset of 4,840 English sentences from financial news, labelled positive, negative, or neutral from an investor's perspective.

#sentiment-analysis #financial-news #classification #annotated #benchmark

added 2026-08-15 · CC-BY-NC-SA-3.0 · external

source English-language financial news
size 4,840 sentences in four agreement-level configurations
coverage assets: 4,840 annotated sentences
access open
license CC-BY-NC-SA-3.0

The standard small-scale benchmark for financial sentiment classification, originating from Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts (JASIST 2014, arXiv:1307.5336). Sentences were annotated by multiple people and the dataset ships in four configurations by inter-annotator agreement: ≥50% (4,846 instances), ≥66% (4,217), ≥75% (3,453), and unanimous (2,264).

Widely used to fine-tune and evaluate financial LLMs. Note the non-commercial share-alike license, and that the Hugging Face dataset viewer does not render it (the repo uses a loading script).

trading [●●···] basic   ai [●●···] basic   programming [●●···] basic   setup [●●···] basic

authors Pekka Malo, Ankur Sinha, Pyry Takala, Pekka Korhonen, Jyrki Wallenius
origin external
markets equities

athenara:~$