datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipediaWikipedia dataset containing cleaned articles of all languages.
The datasets are built from the Wikipedia dump
(https://dumps.wikimedia.org/) with one split per language. Each example
contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).c4A colossal, cleaned version of Common Crawl's web crawl corpus.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's C4 dataset by AllenAI.4chan-datasetsPlease see repo to turn the text file into json/csv format
Deleted some boards, since they are already archived by https://archive.4plebs.org/
llm_datasetstaskmaster2Taskmaster is dataset for goal oriented conversations. The Taskmaster-2 dataset consists of 17,289 dialogs in the seven domains which include restaurants, food ordering, movies, hotels, flights, music and sports. Unlike Taskmaster-1, which includes both written "self-dialogs" and spoken two-person dialogs, Taskmaster-2 consists entirely of spoken two-person dialogs. In addition, while Taskmaster-1 is almost exclusively task-based, Taskmaster-2 contains a good number of search- and recommendation-oriented dialogs. All dialogs in this release were created using a Wizard of Oz (WOz) methodology in which crowdsourced workers played the role of a 'user' and trained call center operators played the role of the 'assistant'. In this way, users were led to believe they were interacting with an automated system that “spoke” using text-to-speech (TTS) even though it was in fact a human behind the scenes. As a result, users could express themselves however they chose in the context of an automated interface.YuLan-Mini-Datasets
YuLan-Mini Datasets
🔥 Updated (April 11, 2025): For a clearer presentation of the information, see the table at this link: link.
This datasets contains:
Classified data using python-edu-scorer and fineweb-edu-classifier
Synthesized data (math, code, instruction, ...)
Retrieved data using math, code, and reasoninig-classifier
Notice
Since we have used BPE-Dropout, in order to ensure accuracy, the data we uploaded is tokenized.… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Datasets.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.amazon_reviews_multiWe provide an Amazon product reviews dataset for multilingual text classification. The dataset contains reviews in English, Japanese, German, French, Chinese and Spanish, collected between November 1, 2015 and November 1, 2019. Each record in the dataset contains the review text, the review title, the star rating, an anonymized reviewer ID, an anonymized product ID and the coarse-grained product category (e.g. ‘books’, ‘appliances’, etc.) The corpus is balanced across stars, so each star rating constitutes 20% of the reviews in each language.
For each language, there are 200,000, 5,000 and 5,000 reviews in the training, development and test sets respectively. The maximum number of reviews per reviewer is 20 and the maximum number of reviews per product is 20. All reviews are truncated after 2,000 characters, and all reviews are at least 20 characters long.
Note that the language of a review does not necessarily match the language of its marketplace (e.g. reviews from amazon.de are primarily written in German, but could also be written in English, etc.). For this reason, we applied a language detection algorithm based on the work in Bojanowski et al. (2017) to determine the language of the review text and we removed reviews that were not written in the expected language.SkillArena-datasets
SkillArena Offline Datasets
Offline evaluation data for SkillArena — a validated automatic benchmark generation framework for AI agent skills, targeting NeurIPS 2026 Datasets & Benchmarks Track.
Overview
This dataset provides domain-specific task input data for 289 AI agent skills across 13 domains. Each skill has 50 curated data files designed as meaningful agent task inputs — files an agent could receive and act upon (analyze, transform, validate, or generate from). The… See the full description on the dataset page: https://huggingface.co/datasets/JiaaqiLiu/SkillArena-datasets.dLLM-PRM-Gap-Datasets
dLLM PRM Gap · Datasets
💻 Code
•
🤗 Collection
•
📄 Paper
This dataset repository contains the selected trajectory corpus and small diagnostic artifacts for dLLM PRM Gap, a controlled study of process reward model (PRM) guidance and outcome reward model (ORM) reranking in discrete diffusion reasoning.
Public release. This dataset accompanies the arXiv paper and the dLLM PRM Gap code release.
News
2026-09: Accepted at NeurIPS… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/dLLM-PRM-Gap-Datasets.MMLA-Datasets
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
1. Introduction
MMLA is the first comprehensive multimodal language analysis benchmark for evaluating foundation models. It has the following features:
Large Scale: 61K+ multimodal samples.
Various Sources: 9 datasets.
Three Modalities: text, video, and audio
Both Acting and Real-world Scenarios: films, TV series, YouTube, Vimeo, Bilibili, TED, improvised scripts, etc.
Six Core… See the full description on the dataset page: https://huggingface.co/datasets/THUIAR/MMLA-Datasets.speculators-ci-datasets
speculator-tutorial
Raw vs. on-policy regenerated conversation data for training speculative-decoding
drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside
so you can see exactly what regeneration changes and why it matters.
Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B.
Why regenerate at all?
A speculative-decoding drafter is trained to predict what the verifier would say next.
If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.schema_guided_dstc8The Schema-Guided Dialogue dataset (SGD) was developed for the Dialogue State Tracking task of the Eights Dialogue Systems Technology Challenge (dstc8).
The SGD dataset consists of over 18k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
These conversations involve interactions with services and APIs spanning 17 domains, ranging from banks and events to media, calendar, travel, and weather.
For most of these domains, the SGD dataset contains multiple different APIs, many of which have overlapping functionalities but different interfaces,
which reflects common real-world scenarios.mc4A colossal, cleaned version of Common Crawl's web crawl corpus.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI.amazon_us_reviewsAmazon Customer Reviews (a.k.a. Product Reviews) is one of Amazons iconic products. In a period of over two decades since the first review in 1995, millions of Amazon customers have contributed over a hundred million reviews to express opinions and describe their experiences regarding products on the Amazon.com website. This makes Amazon Customer Reviews a rich source of information for academic researchers in the fields of Natural Language Processing (NLP), Information Retrieval (IR), and Machine Learning (ML), amongst others. Accordingly, we are releasing this data to further research in multiple disciplines related to understanding customer product experiences. Specifically, this dataset was constructed to represent a sample of customer evaluations and opinions, variation in the perception of a product across geographical regions, and promotional intent or bias in reviews.
Over 130+ million customer reviews are available to researchers as part of this release. The data is available in TSV files in the amazon-reviews-pds S3 bucket in AWS US East Region. Each line in the data files corresponds to an individual review (tab delimited, with no quote and escape characters).
Each Dataset contains the following columns:
- marketplace: 2 letter country code of the marketplace where the review was written.
- customer_id: Random identifier that can be used to aggregate reviews written by a single author.
- review_id: The unique ID of the review.
- product_id: The unique Product ID the review pertains to. In the multilingual dataset the reviews for the same product in different countries can be grouped by the same product_id.
- product_parent: Random identifier that can be used to aggregate reviews for the same product.
- product_title: Title of the product.
- product_category: Broad product category that can be used to group reviews (also used to group the dataset into coherent parts).
- star_rating: The 1-5 star rating of the review.
- helpful_votes: Number of helpful votes.
- total_votes: Number of total votes the review received.
- vine: Review was written as part of the Vine program.
- verified_purchase: The review is on a verified purchase.
- review_headline: The title of the review.
- review_body: The review text.
- review_date: The date the review was written.wiki_snippets
Dataset Card for "wiki_snippets"
Dataset Summary
Wikipedia version split into plain text snippets for dense semantic indexing.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
We show detailed information for 2 configurations of the dataset (with 100 snippet passage length and 0 overlap) in
English:
wiki40b_en_100_0: Wiki-40B
wikipedia_en_100_0: Wikipedia
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/wiki_snippets.to-tool-call-datasets
🛠️ To-Tool-Call Datasets
A unified Qwen3-style tool-call corpus for SFT, GRPO, and agent training
To-Tool-Call Datasets is a curated mirror of public tool-call and function-calling corpora, re-serialized into one training-ready messages JSONL convention.
Quick Start ·
At a Glance ·
Format ·
Sources ·
Training Notes
[!IMPORTANT]
This repository is a format-harmonization layer, not a new claim of ownership over the… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-datasets.YuLan-Mini-Datasets-Phasae-27The tokenized datasets for YuLan-Mini phase 27, where each line has been packed to 28K tokens.
Usage
dataset = []
dataset_path = "/path/to/YuLan-Mini-Datasets-Phasae-27"
seed = 42
for data_name in sorted(os.listdir(dataset_path)):
d = load_dataset(
os.path.join(dataset_path, data_name),
split="train",
num_proc=8,
)
dataset.append(d)
print(f"Num subsets: {len(dataset)}")
dataset = concatenate_datasets(dataset).shuffle(seed=seed)… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Datasets-Phasae-27.YuLan-Mini-Text-Datasets
News
[2025.04.11] Add dataset mixture: link.
[2025.03.30] Text datasets upload finished.
This is text dataset.
这是文本格式的数据集。
Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here.
由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。
For more information, please refer to our datasets details and preprocess details.
Contributing
We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.cpt_instruction_datasets
Instruction datasets
Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3.
Dataset creation
Datasets were created using two different techniques:
Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.YuLan-Mini-Datasets-Phasae-26The tokenized datasets for YuLan-Mini phase 26, where each line has been packed to 28K tokens.
Usage
dataset = []
dataset_path = "/path/to/YuLan-Mini-Datasets-Phasae-26"
seed = 42
for data_name in sorted(os.listdir(dataset_path)):
d = load_dataset(
os.path.join(dataset_path, data_name),
split="train",
num_proc=8,
)
dataset.append(d)
print(f"Num subsets: {len(dataset)}")
dataset = concatenate_datasets(dataset).shuffle(seed=seed)… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Datasets-Phasae-26.AIG-Datasets
AMD AIG GPU Kernel Datasets
AMD AIG-Datasets is a collection of GPU-kernel generation, translation, optimization,
and ROCm-library supervision data. It contains PyTorch/CUDA-to-HIP, HIP-to-HIP,
PyTorch-to-Triton, and production-grounded rocBLAS/rocSOLVER entries, together with
metadata, samples, conversion utilities, and reproducible evaluation tools.
The repository is organized into versioned releases. New training and evaluation
workflows should use the unified-schema datasets… See the full description on the dataset page: https://huggingface.co/datasets/amd/AIG-Datasets.the_pile_books3This dataset is Shawn Presser's work and is part of EleutherAi/The Pile dataset. This dataset contains all of bibliotik in plain .txt form, aka 197,000 books processed in exactly the same way as did for bookcorpusopen (a.k.a. books1). seems to be similar to OpenAI's mysterious "books2" dataset referenced in their papers. Unfortunately OpenAI will not give details, so we know very little about any differences. People suspect it's "all of libgen", but it's purely conjecture.fractus-datasets
Fractus Datasets — the neuroscience-grounded training corpus
A proprietary, neuroscience-derived training corpus for the Fractus Continuous Thought Engine — ~3–4B tokens mapping real brain mechanisms to software/AI architecture, plus cognitive skills, code, esoteric tradition, and lexical knowledge.
Curator: Philippe-Antoine Robert · rpa.tu@proton.me · 2026
What this dataset collection IS
Fractus is a non-transformer Continuous Cognitive Agent whose architecture… See the full description on the dataset page: https://huggingface.co/datasets/thefinalboss/fractus-datasets.SLCA-GRPO-Datasets
SLCA-GRPO · Datasets
📄 Paper (arXiv:2609.29050)
•
💻 Code
•
🤗 Collection
This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit
assignment estimator for tool-calling RL introduced in "SLCA-GRPO: Resolving Cross-Segment Credit
Misattribution in Tool-Calling RL". A tool-calling rollout opens with structured tool-call tokens
and closes with free-form summary text; SLCA-GRPO normalises the two segment rewards… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/SLCA-GRPO-Datasets.bookcorpusopenBooks are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story.
This version of bookcorpus has 17868 dataset items (books). Each item contains two fields: title and text. The title is the name of the book (just the file name) while text contains unprocessed book text. The bookcorpus has been prepared by Shawn Presser and is generously hosted by The-Eye. The-Eye is a non-profit, community driven platform dedicated to the archiving and long-term preservation of any and all data including but by no means limited to... websites, books, games, software, video, audio, other digital-obscura and ideas.chess_datasets
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/p11-p11/chess_datasets.ToolGen-Datasets
How to use?
Before making use of this dataset, you may need to add the tokens to the vocabulary. For HuggingFace transformers tokenizer, the following is an example code snippet to add tokens.
from unidecode import unidecode
import transformers
with open('virtual_tokens.txt', 'r') as f:
virtual_tokens = f.readlines()
virtual_tokens = [unidecode(vt.strip()) for vt in virtual_tokens]
model_name_or_path = "meta-llama/Meta-Llama-3-8B"
# Load tokenizer and add tokens into… See the full description on the dataset page: https://huggingface.co/datasets/reasonwang/ToolGen-Datasets.capstone-datasets
AIMS AI Research Foundations – Capstone Datasets
Ready-to-use train / validation / test datasets for the AIRF Capstone Project Recipes, a set of short, end-to-end recipes for implementing and evaluating AI capstone projects, for university lecturers and learners across Africa, built around small open-weight models (Gemma 1B / 4B).
Every dataset is a configuration of this repository. Every configuration has train, validation and test splits with labels or reference answers.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/capstone-datasets.SafeCRS-datasets
SafeCRS Datasets
Datasets for training safety-aware conversational recommendation models. The
default viewer shows the SafeRec SFT test split — the safety-annotated
evaluation set with one explicit user sensitivity trait per sample.
SafeRec SFT Dataset (test split preview)
Each sample contains a user conversation requesting movie recommendations, an
assigned sensitivity trait, constraint-filtered ground truth, and the
constraint text injected into the prompt.… See the full description on the dataset page: https://huggingface.co/datasets/Dionysianspirit/SafeCRS-datasets.
