datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finance_agent_benchmark
Finance Agent Benchmark Dataset
We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings.
We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.FinanceQAFinanceQA is a comprehensive testing suite designed to evaluate LLMs' performance on complex financial analysis tasks that mirror real-world investment work. The dataset aims to be substantially more challenging and practical than existing financial benchmarks, focusing on tasks that require precise calculations and professional judgment.
Paper: https://arxiv.org/abs/2501.18062
Description
The dataset contains two main categories of questions:
Tactical Questions: Questions based on… See the full description on the dataset page: https://huggingface.co/datasets/AfterQuery/FinanceQA.Sujet-Finance-Instruct-177k
Sujet Finance Dataset Overview
The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.XBRL_financebench_metricsFinanceIQus-bank-insurance-finance-layoffs-warn-act-notices-daily
US bank, insurance and finance layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-24. 2,101 layoff and closure notices filed by
banks and mortgage lenders, insurers and health-plan administrators, brokerages and asset managers, card and payments companies, credit unions, tax, payroll and accounting firms, and real-estate and property-management operators with US state labor departments — 180,918 workers,
768 employers, 45 states, 1988–2026.
405 of the… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-bank-insurance-finance-layoffs-warn-act-notices-daily.Finance-Competitiveness-and-Innovation-Indicators-For-African-Countries
Finance Competitiveness and Innovation Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Finance-Competitiveness-and-Innovation-Indicators-For-African-Countries.trending-stocks-yahoo-finance
Monthly Trending Stocks Dataset
A ranked dataset of the most trending stocks on Yahoo Finance from July 2024 to October 2025, based on weighted scoring of their monthly trending appearances in Yahoo Finance.
📊 Dataset Overview
Total Entries: 7,993 ranked stocks
Time Period: July 2024 - October 2025 (16 months)
Source: Wayback Machine snapshots of Yahoo Finance Trending Stocks
Data Granularity: Monthly rankings
Data Order: Sorted by month (descending: Oct 2025 → July… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/trending-stocks-yahoo-finance.finance-news-sentiment-35k
Finance News Sentiment 40k
39,965 English financial news headlines, collected from public Telegram
finance news-wire channels, labeled for 3-class sentiment (positive / negative / neutral) and a secondary
topic label, by two independent LLM judges from different model families with
an arbiter settling disputes.
A FinBERT model fine-tuned on this data reaches test accuracy 0.847 / macro F1
0.810: remehostingservices/finbert-finance-news-sentiment.
Code, training scripts and the… See the full description on the dataset page: https://huggingface.co/datasets/remehostingservices/finance-news-sentiment-35k.reddit_finance_posts_sp500
Reddit Finance Posts Dataset SP500
This dataset contains 431,923 Reddit posts collected from 20 finance-related subreddits via the Reddit API.
As keywords, all S&P 500 companies were used. The included subreddits are:
stocks, wallstreetbets, investing, StockMarket, options, RobinHood,
pennystocks, SecurityAnalysis, personalfinance, Dividends, CryptoCurrency,
CryptoMarkets, ETFs, FinancialIndependence, ValueInvesting, quant,
algotrading, forex, economy, Superstonk, spacs… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/reddit_finance_posts_sp500.embedding-finetuning-financeThis dataset can be used for fine-tuning embedding models using positive text pairs (question, context).
finance-datasets
Finance Datasets
Historical stock and cryptocurrency price data.
Contents
Stocks (5 years of daily OHLCV data)
AAPL - Apple Inc.
GOOGL - Alphabet Inc.
MSFT - Microsoft Corp.
AMZN - Amazon.com Inc.
TSLA - Tesla Inc.
META - Meta Platforms
NVDA - NVIDIA Corp.
AMD - Advanced Micro Devices
INTC - Intel Corp.
NFLX - Netflix Inc.
Cryptocurrencies (full history)
BTC_USD - Bitcoin
ETH_USD - Ethereum
SOL_USD - Solana
ADA_USD - Cardano
DOT_USD - Polkadot… See the full description on the dataset page: https://huggingface.co/datasets/misterdonn/finance-datasets.nlp-finance-project1
Project 1 — News Headlines: Sentiment Lexicon vs. Instructor Baseline
Warm-up project for the final NLP-for-Finance project. Builds on HW1 (lexicon-based
sentiment) and HW2 (topic modelling) using the same S&P/DOW/NASDAQ news-headline
dataset (wiliamvvvvv/SP_DOW_NASDAQ_stocks__News_Headlines_labeled).
1. Data
91,851 ticker-day rows: 84,643 train / 4,811 test / 2,397 inference (instructor's
original date-based split, revision lesson1_sentiment_analysis).
Target:… See the full description on the dataset page: https://huggingface.co/datasets/wiliamvvvvv/nlp-finance-project1.finance_agent_benchmark
Finance Agent Benchmark Dataset
We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings.
We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/finance_agent_benchmark.denial-ai-benchmark
The Denial-AI Benchmark (v1.3, frozen)
Twelve fixed questions about US FHA mortgage denial outcomes, each with a ground truth computed from the complete 2025 federal HMDA record
(FHA credit decisions, reverse mortgages excluded; universe and filters at https://financeratecalc.com/methodology.html)
and a source where the figure is published and reproducible.
The questions are frozen; re-administration measures improvement or drift. Instrument license: CC BY 4.0.
Latest… See the full description on the dataset page: https://huggingface.co/datasets/FinanceRateCalc/denial-ai-benchmark.Finance-Text-Clasification
Finance Text Classification Dataset
A lightweight English-language dataset designed for text classification tasks in the finance domain.It contains 1,000 synthetic but realistic loan application texts with requested amounts up to €50,000.
The dataset is suitable for:
Binary text classification (approved vs. rejected)
Zero-shot classification experiments
Fine-tuning or evaluating language models on financial application texts
Dataset Description
Each example… See the full description on the dataset page: https://huggingface.co/datasets/OceanLabs/Finance-Text-Clasification.vidore3_finance_neomme_260m_li
vidore3_finance_neomme_260m_li
Multi-vector (late-interaction) embeddings of ViDoRe finance (vidore/finance), encoded with
Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.
Source data: Hugging Face dataset vidore/vidore_v3_finance_en at revision 7f432c176d82e27546501ad8064a713ac3071809, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids,
unchanged.
Every document is… See the full description on the dataset page: https://huggingface.co/datasets/robro612/vidore3_finance_neomme_260m_li.FinanceQAFinanceQA is a comprehensive testing suite designed to evaluate LLMs' performance on complex financial analysis tasks that mirror real-world investment work. The dataset aims to be substantially more challenging and practical than existing financial benchmarks, focusing on tasks that require precise calculations and professional judgment.
Paper: https://arxiv.org/abs/2501.18062
Description
The dataset contains two main categories of questions:
Tactical Questions: Questions based on… See the full description on the dataset page: https://huggingface.co/datasets/Joshua-Xia/FinanceQA.accounting-finance-audit-open-registry
Accounting, Finance & Audit Open Data and Program Registry
Cross-platform public registry maintained from the canonical GitHub repository of Saeid Homayoun.
It connects free/public resources across GitHub, Hugging Face and Kaggle for accounting, auditing, assurance, finance, fraud/AML, SEC/EDGAR/XBRL, ESG/sustainability, Microsoft/Copilot and IBM/watsonx research.
What this mirror contains
program_catalog.csv — curated public/free program catalogue.… See the full description on the dataset page: https://huggingface.co/datasets/SADHON/accounting-finance-audit-open-registry.vidore3_financefr_neomme_260m_li
vidore3_financefr_neomme_260m_li
Multi-vector (late-interaction) embeddings of ViDoRe financefr (vidore/financefr), encoded with
Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.
Source data: Hugging Face dataset vidore/vidore_v3_finance_fr at revision 1d808daa08032ffecdf62da151a7f7a8fe2bd0c9, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids,
unchanged.
Every… See the full description on the dataset page: https://huggingface.co/datasets/robro612/vidore3_financefr_neomme_260m_li.Financial-Reportsachintyatripathi_yahoo-finance-apple-inc-aapl
Yahoo Finance Apple Inc. (AAPL)
Try Predicting the market
Dataset Info
Source: Kaggle
Original Size: 0.01 MB
Kaggle Downloads: 3,990
Files: 3
Files
AAPL_Montly_updates.csv
AAPL_daily_update.csv
AAPL_weekly_update.csv
Mirrored from Kaggle
indic-finance
🇮🇳 Indic-Finance-Sentiment-Hub
A high-frequency, multi-source dataset of Indian financial news, social media discussions, and stock price movements. This dataset is designed for training sentiment analysis models (like FinBERT) and predictive models for the Indian Equity Market (NSE/BSE).
📊 Dataset Summary
The dataset aggregates headlines and posts from 150+ top Indian companies (Nifty 50, Nifty Next 50, and major mid-caps). Each observation is labeled with a FinBERT… See the full description on the dataset page: https://huggingface.co/datasets/dixitdharmansh07/indic-finance.aws-free-accounting-audit-finance-registry
AWS Free Data Registry for Accounting, Audit, Finance & ESG
A provenance-first registry curated for NAAIL OpenLab™ / FRANKENSTEIN™.
It identifies free/open AWS-hosted or AWS-delivered sources useful for accounting, auditing, assurance, finance, macroeconomic analysis, ESG, carbon accounting, Scope 3 and climate-risk research.
Important
This dataset mirrors metadata and links only. It does not redistribute third-party AWS Marketplace datasets whose vendor terms may… See the full description on the dataset page: https://huggingface.co/datasets/SADHON/aws-free-accounting-audit-finance-registry.predator-finance
PREDATOR Finance Dataset
700+ curated finance reports from DataGov and public sources
Financial data including market reports, earnings summaries, and regulatory filings. Useful for training financial NLP models, sentiment analysis, and market research agents.
Fields
Column
Description
id
Report identifier
title
Report title or summary
source
Data source (DataGov, SEC, public filings)
domain
Classification (finance, earnings, regulatory)… See the full description on the dataset page: https://huggingface.co/datasets/iservice/predator-finance.financebench_pdftiny-aya-global-finance-evalreddit_finance_posts_apple-tesla-microsoft
Reddit Finance Posts Dataset (Apple, Tesla, Microsoft)
This dataset contains 12046 Reddit Posts collected from 20 finance-related subreddits via the Reddit API using the keywords Apple, Tesla, and Microsoft.
The included subreddits are:
stocks, wallstreetbets, investing, StockMarket, options, RobinHood,
pennystocks, SecurityAnalysis, personalfinance, Dividends, CryptoCurrency,
CryptoMarkets, ETFs, FinancialIndependence, ValueInvesting, quant,
algotrading, forex, economy, Superstonk… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/reddit_finance_posts_apple-tesla-microsoft.finance_sentimentfinguard-finance-injection-dataset
FinGuard: Finance-Specific Prompt Injection Detection Dataset
Dataset Summary
FinGuard is the first open dataset for detecting prompt injection attacks
against agentic financial AI systems. It combines 6 public datasets with
synthetically generated finance-specific attack examples across 4 enterprise
agent types.
Dataset Structure
Split
Rows
SAFE
ATTACK
Train
10,699
5,375 (50.2%)
5,324 (49.8%)
Test
3,047
2,006 (65.8%)
1,041 (34.2%)… See the full description on the dataset page: https://huggingface.co/datasets/nandhak12/finguard-finance-injection-dataset.
