Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01defeatbeta /yahoo-finance-data The Financial data from Yahoo! *** Key Points to Note *** All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes. I will update the data regularly, and you are welcome to follow this project and use the data. Each time the data is updated, I will record the update time in spec.json. Data Usage Instructions Use DuckDB… See the full description on the dataset page: https://huggingface.co/datasets/defeatbeta/yahoo-finance-data.100M<n<1B133 likes129k downloads6h agoHugging Face02PatronusAI /financebenchFinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). This is an open source sample of 150 annotated examples used in the evaluation and analysis of models assessed in the FinanceBench paper. The PDFs linked in the dataset can be found here as well: https://github.com/patronus-ai/financebench/tree/main/pdfs The dataset comprises of questions about publicly traded companies, with corresponding answers and evidence… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/financebench.textn<1K142 likes6.8k downloads2y agoHugging Face03dragonlimited /DragonData-Finance-Corpus DragonData World Finance Database Priced like infrastructure. Free like air. The institutional-grade bilingual (EN/中文) finance database for training and evaluating expert AI systems — collected, cleaned, deduplicated, tokenized, license-proven, timestamped. 3,700+ files and counting. Start here: Browse the data ⬇ · Use it in 4 lines ⬇ · Live build feed ⬇ Browse the data The table below is live — scroll, search, and read actual records right on this page.… See the full description on the dataset page: https://huggingface.co/datasets/dragonlimited/DragonData-Finance-Corpus.texttext-generation1K<n<10K0 likes4.1k downloads22h agoHugging Face04vidore /vidore_v3_finance_enViDoRe V3 : Finance - EN This dataset, Financial_Bank_Reports, is a corpus of annual reports from the banking sector, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_en.documentvisual-document-retrieval10K<n<100K17 likes2.7k downloads9mo agoHugging Face05nvidia /Nemotron-SpecializedDomains-Finance-v1 Dataset Description Nemotron-SpecializedDomains-Finance is a large-scale synthetic financial question-answering dataset designed to improve LLM performance on specialized financial reasoning and document comprehension tasks. The dataset comprises 326K+ high-quality Q&A pairs generated from SEC filings of S&P 500 companies spanning 2019-2024. This dataset is ready for commercial use. Overview The dataset leverages template-based Synthetic Data Generation (SDG) to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1.texttext-generation100K<n<1M16 likes2.7k downloads7mo agoHugging Face06BAAI /IndustryCorpus_finance[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_finance.texttext-generation10M<n<100M19 likes2.6k downloads2mo agoHugging Face07gbharti /finance-alpacaThis dataset is a combination of Stanford's Alpaca (https://github.com/tatsu-lab/stanford_alpaca) and FiQA (https://sites.google.com/view/fiqa/) with another 1.3k pairs custom generated using GPT3.5 Script for tuning through Kaggle's (https://www.kaggle.com) free resources using PEFT/LoRa: https://www.kaggle.com/code/gbhacker23/wealth-alpaca-lora GitHub repo with performance analyses, training and data generation scripts, and inference notebooks: https://github.com/gaurangbharti1/wealth-alpaca… See the full description on the dataset page: https://huggingface.co/datasets/gbharti/finance-alpaca.texttext-generation10K<n<100K156 likes2.6k downloads11mo agoHugging Face08gretelai /synthetic_pii_finance_multilingual Image generated by DALL-E. See prompt for more details 💼 📊 Synthetic Financial Domain Documents with PII Labels gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0. This dataset is designed to assist with the following use cases: 🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.tabulartext-classification10K<n<100K81 likes2.4k downloads2y agoHugging Face09AdaptLLM /finance-tasks Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/finance-tasks.tabulartext-classification10K<n<100K83 likes2.1k downloads2y agoHugging Face10BEE-spoke-data /consumer-finance-complaints BEE-spoke-data/consumer-finance-complaints consumer-finance-complaints but in a format that actually works. Pulled Feb 2024 texttext-classification1M<n<10M5 likes2k downloads10mo agoHugging Face11vals-ai /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.textn<1K11 likes1.9k downloads1y agoHugging Face12noetic-labs /finance-data-gdpval finance-data — DHR / ILMN M&A Analysis Tasks Investment-banking analysis tasks set in a (fictional) strategic M&A project where Danaher (DHR) explores acquiring Illumina (ILMN). Each task gives an agent a prompt plus a full deal-room of reference materials (financial models, spreadsheets, PDFs, research) and grades the answer against a per-criterion rubric. Tasks that require modifying a workbook also ship the expert's gold workbook as a deliverable file. Layout… See the full description on the dataset page: https://huggingface.co/datasets/noetic-labs/finance-data-gdpval.documentn<1K0 likes1.8k downloads3mo agoHugging Face13artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes1.8k downloads8mo agoHugging Face14Joshua-Xia /yahoo-finance-data The Financial data from Yahoo! *** Key Points to Note *** All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes. I will update the data regularly, and you are welcome to follow this project and use the data. Each time the data is updated, I will record the update time in spec.json. Data Usage Instructions Use DuckDB or… See the full description on the dataset page: https://huggingface.co/datasets/Joshua-Xia/yahoo-finance-data.100M<n<1B0 likes1.5k downloads8mo agoHugging Face15vidore /vidore_v3_finance_frViDoRe V3 : Finance - FR This dataset, Finance - FR, is a corpus of reports from french companies in the luxury domain, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_fr.documentvisual-document-retrieval10K<n<100K3 likes1.4k downloads9mo agoHugging Face16vidore /vidore_v3_finance_en_mteb_format Vidore3FinanceEnRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_finance_en How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3FinanceEnRetrieval") evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_en_mteb_format.imagevisual-document-retrieval10K<n<100K1 likes1.4k downloads11mo agoHugging Face17embedding-benchmark /FinanceBenchThe FinanceBench dataset is derived from the PatronusAI/financebench-test dataset, containing only the PASS examples processed into a clean format for question-answering tasks in the financial domain. FinanceBench-rtl has been repurposed for retrieval. Usage import datasets # Download the dataset queries = datasets.load_dataset("embedding-benchmark/FinanceBench", "queries") documents = datasets.load_dataset("embedding-benchmark/FinanceBench", "corpus") pair_labels =… See the full description on the dataset page: https://huggingface.co/datasets/embedding-benchmark/FinanceBench.texttext-retrievaln<1K0 likes1.2k downloads1y agoHugging Face18GGLabYale /MTBench_finance_stock MTBench: A Multimodal Time Series Benchmark MTBench (Huggingface, Github, Arxiv) is a suite of multimodal datasets for evaluating large language models (LLMs) in temporal and cross-modal reasoning tasks across finance and weather domains. Each benchmark instance aligns high-resolution time series (e.g., stock prices, weather data) with textual context (e.g., news articles, QA prompts), enabling research into temporally grounded and multimodal understanding. 🏦 Stock… See the full description on the dataset page: https://huggingface.co/datasets/GGLabYale/MTBench_finance_stock.1K<n<10K2 likes1.2k downloads1y agoHugging Face19thomaskim1130 /FinanceRAG-Linguatext100K<n<1M2 likes1.1k downloads2y agoHugging Face20financeindustryknowledgeskills /modeling_valuation_knowledge Finance Training Data Repository A curated collection of financial modeling courses, materials, and resources designed to serve as training data for building a finance industry knowledge base. Repository Structure Finance_Training_Data/ ├── 01_Financial_Statement_Modeling/ # 3-statement modeling fundamentals ├── 02_DCF_Modeling/ # Discounted cash flow valuation ├── 03_Trading_Comps/ # Comparable company analysis ├──… See the full description on the dataset page: https://huggingface.co/datasets/financeindustryknowledgeskills/modeling_valuation_knowledge.documentn<1K0 likes1.1k downloads4mo agoHugging Face21Yahoo-Finance-News /FineWeb2024 FineWeb-Edu 2024 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2024. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2024 Rows 162,500,784… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2024.tabulartext-generation100M<n<1B0 likes1.1k downloads22d agoHugging Face22RogoAI /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.textquestion-answeringn<1K13 likes1k downloads2mo agoHugging Face23davidheineman /consumer-finance-complaints-large CFPB Complaints A dataset of 7M complaints from the Consumer Financial Protection Bureau (CFPB), from 12/01/2011 to 01/02/2025. For descriptions of each column, please see consumerfinance.gov/complaint/data-use. More Details A description from the CFPB website: The Consumer Complaint Database is a collection of complaints about consumer financial products and services that we sent to companies for response. Complaints are published after the company responds, confirming… See the full description on the dataset page: https://huggingface.co/datasets/davidheineman/consumer-finance-complaints-large.text1M<n<10M1 likes1k downloads2y agoHugging Face24AfterQuery /FinanceQAFinanceQA is a comprehensive testing suite designed to evaluate LLMs' performance on complex financial analysis tasks that mirror real-world investment work. The dataset aims to be substantially more challenging and practical than existing financial benchmarks, focusing on tasks that require precise calculations and professional judgment. Paper: https://arxiv.org/abs/2501.18062 Description The dataset contains two main categories of questions: Tactical Questions: Questions based on… See the full description on the dataset page: https://huggingface.co/datasets/AfterQuery/FinanceQA.textquestion-answeringn<1K24 likes991 downloads2y agoHugging Face25Yahoo-Finance-News /FineWeb2025 FineWeb-Edu 2025 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2025 Rows 99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.tabulartext-generation10M<n<100M1 likes948 downloads22d agoHugging Face26FinanceMTEB /FinQAtabular10K<n<100K1 likes941 downloads2y agoHugging Face27oss-codes /Finance-Conversational-Dataset-Indictext100K<n<1M1 likes913 downloads2y agoHugging Face28BAAI /IndustryCorpus2_finance_economics IndustryCorpus2: Finance & Economics This repository contains the IndustryCorpus2: Finance & Economics domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_finance_economics.tabular10M<n<100M11 likes907 downloads2mo agoHugging Face29FinanceMTEB /TATQAtabular1K<n<10K0 likes875 downloads2y agoHugging Face30handshake-ai-research /ATLAS-Finance ATLAS Finance A benchmark of 100 expert-level tasks inside 13 realistic financial firm environments, packaged in the Harbor RLE format. Each task drops an AI agent into a Linux workstation with a persistent multi-app world — inbox, chat, calendar, virtual data room, drive, wiki — and asks the agent to produce the same deliverable a financial professional would be responsible for: an Excel workbook containing the model and supporting analysis. Here we provide the data for this… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/ATLAS-Finance.documenttext-generationn<1K4 likes875 downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.