datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipo-text
SEC IPO Filings Dataset
A large-scale, comprehensive dataset of 100,000+ filings (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026 and over 20,000 unique registrants.
Every filing has been downloaded and then parsed using the IPO-Mine Python Package. We have extracted three common sections found in these documents (Prospectus Summary, Risk Factors, Legal Matters), and then used an LLM classifier to group them into three categories. For this dataset, we have… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-text.persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.WTO-Text
Dataset Card for WTO Documents Dataset
Dataset Overview
Title: WTO Documents Dataset
Source: World Trade Organization Documents Online
Description: The WTO Documents Dataset is a comprehensive collection of official documentation from the World Trade Organization (WTO). This dataset is sourced from the WTO's official Documents Online platform, which provides access to documents in the three official languages (English, French, and Spanish) from 1995 onwards. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/WTO-Text.TextOnly_FromRLBench_CloseBox_24K_unfixedArgo-Bench
Argo-Bench
One simulated year of a New York City food-delivery company, exported to an Oracle E-Business
Suite 12.2 warehouse of 235 tables and 7.54 billion rows.
Leaderboard & Demo · How it works ·
Paper · Code
An enterprise ERP you can download
The data that enterprise analytics runs on is rarely public. Large companies keep their orders,
payouts, ledgers and customer records in ERP systems such as Oracle E-Business Suite and SAP,
extended with custom tables of… See the full description on the dataset page: https://huggingface.co/datasets/textql/Argo-Bench.Decision-Bench
Decision-Bench
One simulated year of a New York City food-delivery company, exported to an Oracle E-Business
Suite 12.2 warehouse of 235 tables and 7.54 billion rows.
Leaderboard & Demo · How it works ·
Paper · Code
An enterprise ERP you can download
The data that enterprise analytics runs on is rarely public. Large companies keep their orders,
payouts, ledgers and customer records in ERP systems such as Oracle E-Business Suite and SAP,
extended with custom… See the full description on the dataset page: https://huggingface.co/datasets/textql/Decision-Bench.tiny-aya-global-em-en-text-insecuretext-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.text-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.SQaLe-text-to-SQL-dataset
🧮 SQALE: A Large-Scale Semi-Synthetic Dataset
SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas.
It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions.
The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.python-text-copilot-training-instruct-ai-research-2024-02-03
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab:
Agora GitHub Organization
Agora Hugging Face
This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.text-2-video-human-preferences-veo3
Rapidata Video Generation Veo 3 Human Preference
In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.text_ratingsTodo - Write dataset card
Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.text-code-galeras-code-generation-from-docstring-3k-dedupedtiny-textbooks
Textbook-like Dataset: A High-Quality Resource for Small Language Models
The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model.
Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.EEG-semantic-text-relevanceWe release a novel dataset containing 23,270 time-locked (0.7s) word-level EEG recordings acquired from participants who read both text that was semantically relevant and irrelevant to self-selected topics.
The raw EEG data and the datasheet are available at https://osf.io/xh3g5/.
See code repository for benchmark results.
EEG data acquisition:
Explanations of the variables:
event corresponds to a specific point in time during EEG data collection and represents the onset of an event… See the full description on the dataset page: https://huggingface.co/datasets/Quoron/EEG-semantic-text-relevance.YuLan-Mini-Text-Datasets
News
[2025.04.11] Add dataset mixture: link.
[2025.03.30] Text datasets upload finished.
This is text dataset.
这是文本格式的数据集。
Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here.
由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。
For more information, please refer to our datasets details and preprocess details.
Contributing
We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.text-to-sql-shop
Text-to-SQL on a seeded store schema, with checkpoints
Recipe: recipes/04-train/text-to-sql · Collection: Analyst
A question about an online store's database in, one PostgreSQL query out, graded by
a program: run the query, compare the result set to the gold query's result. The
schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and
the benchmark runner are the
recipes/04-train/text-to-sql
recipe in the open-source whileai SDK.
Splits
|… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.icl-selfplay-vs-text-results
Self-play vs. text pretraining: ICL results
These are the full evaluation results from comparing the in-context learning (ICL) of:
the self-play learners from
Self-Play Pretraining with Zero Data, and
same-size models trained on ordinary web text for the same number of tokens.
Code and write-up: github.com/mihir-s-05/icl-selfplay-vs-text.
Text-model checkpoints: rihim/icl-selfplay-vs-text-checkpoints.
What was scored
Model families (the arm column):
arm… See the full description on the dataset page: https://huggingface.co/datasets/rihim/icl-selfplay-vs-text-results.text-2-video-human-preferences-moonvalley-marey
Rapidata Video Generation Marey Pro Human Preference
In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.text-2-video-human-preferences-veo3.1
Rapidata Video Generation Veo 3.1 Human Preference
In this dataset, ~74k human responses from ~23k human annotators were collected to evaluate the Veo 3.1 video generation model on our benchmark. This dataset was collected using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it ❤️… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.1.python-text-copilot-training-instruct
Python Copilot Instructions on How to Code using Alpaca and Yaml
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.spider-text-to-sql
Spider Text-to-SQL with LLM-Judge Labels
This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4.
Files
File
Description
spider_dataset.parquet
Full dataset with predictions and labels
scripts/
Reproduction scripts (see below)
Dataset statistics
Source: Spider 1.0 training split (train_spider.json)
Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5
FineWeb-edu 10BT Sample embedded with nomic-text-v1.5
The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT.
The chunks were then embedded using nomic-text-v1.5.
Dataset Details
Dataset Sources
Repository: https://github.com/enjalot/fineweb-modal
Uses
Direct Use
The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.python-text-copilot-training-instruct-ai-research-2024-02-11
Python Copilot Instructions on How to Code using Alpaca and Yaml
Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project:
Qwen
Qwen Agent
Qwen VL Chat
Qwen Audio
This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.datacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
alania-domain-text-tr
Alania Turkish Domain Text
English · Türkçe
1,439,639 unique Turkish sentences of the kind a voice agent actually says: appointment and clinic dialogue,
banking, e-commerce and cargo, customer support, confirmations, questions, numbers and codes, addresses and
names, empathetic lines and longer explanations. About 1,819 hours of speech if read aloud.
We built it as the text side of synthetic training speech for Alania-2, the Turkish text-to-speech model behind… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr.text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.cogito-text-world
Cogito Text World
Part of Cogito Capability Checks — does your model actually use what it read?
Small synthetic tests with exact answers, guessing floors and a length ladder (1k → 128k tokens).
Project page: ai.ksopyla.com/projects/cogito-capability-datasets
TL;DR. Train a small language model from scratch on simple English stories with facts woven in,
then ask it about those facts in documents of 1k to 128k tokens (trained at 4k). Seven question
types, from copying a sign to… See the full description on the dataset page: https://huggingface.co/datasets/ksopyla/cogito-text-world.
