datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.fine-news-sample
Fine-News Sample
Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus.
The sample covers all 117 capture months and 388 language-and-script labels in that corpus.
Each selected row preserves its article text, source metadata, and sampling weight.
The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives.
At a glance
Measure
Value
Rows
1,000,000
Distinct document IDs
1,000,000
Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.dolma3_300B_sample
Dolma 3 — 300B-token sample
🌐 The Fin AI
Pretraining / reference corpus released by The Fin AI. Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai.
Source
A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0.
Structure
Rows: 187,823,645
Columns: source, date, text, token_count, category
Quick Start
from datasets import load_dataset
ds = load_dataset("TheFinAI/dolma3_300B_sample"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample.agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.egocentric-activity-sample
Egocentric Activity Sample Dataset
A small-scale egocentric (first-person) video dataset with Ego4D-style annotations, designed for quick prototyping and experimentation with egocentric video understanding tasks.
Dataset Summary
Metric
Value
Video clips
19
Total duration
~9.5 minutes
Resolution
960x540 (540p)
FPS
30
Narrations
99
NLQ queries
57
Moment annotations
19
FHO actions
57
Total size
~54 MB
Activities Covered… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/egocentric-activity-sample.fineweb-edu-sample-10BT-shuffled
📚 FineWeb-Edu (Shuffled)
The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves.
This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu.
Shuffling was performed using the following script:
import datasets
data = datasets.load_dataset(
"HuggingFaceFW/fineweb-edu",
"sample-10BT",
split="train",
streaming=False,
)
data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.reddit-comments-sample
Reddit Comment Trees Sample — Initial snapshot
Initial sample: 43,913 posts and 59,874 comments across three communities.
This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive.
Overview
Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction.
The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.dolma3_300B_sample_shuffled
dolma3_300B_sample_shuffled
🌐 The Fin AI
Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai. A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0. License: ODC BY.
Global row-level shuffle of TheFinAI/dolma3_300B_sample.
Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from
allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving
the original Dolma3 mix ratios. However the source parquets cluster
records by… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample_shuffled.m-a-p-FineFineWeb-sample
Unofficial m-a-p/FineFineWeb Sample
This dataset is a processed, lightweight sample of the original m-a-p/FineFineWeb, a comprehensive corpus designed for fine-grained domain web text studies.
Sampling Methodology
To create this subset, the following processing steps were taken:
Selection: 100 random .jsonl files were chosen from the original dataset.
Extraction: 10,000 rows were downloaded per selected file.
Processing: The extracted rows were combined and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/m-a-p-FineFineWeb-sample.thomas-yanxin-MT-SFT-ShareGPT-sample
MT-SFT-ShareGPT Sample Dataset
This dataset provides a sample of the thomas-yanxin/MT-SFT-ShareGPT dataset with English and Chinese subsets.
Dataset Contents
train.jsonl: Contains 1/10 of the original data, shuffled
EN.jsonl: English conversations from train.jsonl
ZH.jsonl: Chinese conversations from train.jsonl
Each row represents a conversation with an optional system message, followed by human and GPT turns.
Columns from the original dataset are preserved, with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/thomas-yanxin-MT-SFT-ShareGPT-sample.Usenet-Corpus-1980-2013-Threaded-Samples
Usenet Corpus 1980–2013 — Threaded (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded
dataset: Usenet posts reconstructed into conversations via thread_id,
thread_position, and thread_depth. This repo is a free preview; the full,
commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at:
Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded
Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.fineweb-edu-2016-qwen2-sample
FineWeb-Edu 2016 / Qwen2
Consistency sample — not the completed year.
Documents: 900. Actual recounted Qwen2 tokens: 937,977.
Source: HuggingFaceFW/fineweb-edu, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9.
The input inventory covers 9 crawl directories. date is the integer crawl year 2016,
not an article publication date. Original text is preserved without cleaning,
normalization, deduplication, truncation, or added formatting. Source token counts
are not used. Optional… See the full description on the dataset page: https://huggingface.co/datasets/BoomQ/fineweb-edu-2016-qwen2-sample.HuggingFaceFW-finewiki-sample
HuggingFaceFW/finewiki sample
A uniformly randomized subset of HuggingFaceFW/finewiki, created to provide a smaller and more manageable dataset for analysis, fine-tuning, and benchmarking.
Overview
This sample includes Wikipedia articles from languages with more than one million pages. Sampling is performed uniformly at random instead of alphabetically to ensure unbiased representation.
Language Inclusion Criteria
Languages were selected based on page count and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finewiki-sample.common-crawl-docx-sample
Common Crawl DOCX Sample
A sample of normalized text extracted from DOCX records in Common Crawl.
Source
Common Crawl release: CC-MAIN-YYYY-NN
Source index: Common Crawl URL Index
Pipeline: marin-community/marin
Pipeline revision: REPLACE_WITH_GIT_SHA
Records were selected using declared DOCX MIME type, detected DOCX MIME type,
or a .docx URL suffix. Only successful, non-truncated index records were
eligible.
Processing
The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.TACTBench-Samples
TACTBench Demonstration Samples
This repository contains five full-context demonstration examples from
TACTBench. It does not contain the TACT training set or the remaining hidden
TACTBench evaluation set. The samples use the same full-history representation
as the benchmark evaluation and illustrate direct correction, error
explanation, guided revision, clarification checking, affective feedback, and
retry elicitation.
Data
data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.french-corpus-llm-sample
French Corpus LLM — Sample 500 (v1.4.0)
FINALEADS LLC builds compliance-ready training datasets for French regulated industries. We turn 2.66 billion tokens of French finance, regulatory, and economic open data into audit-trailed, pseudonymized, AI Act Article 10-documented shares — so foundation model and regtech teams can ship into European enterprises without a data-lineage gap.
This is a public sample of 500 stratified documents drawn from the French Premium Web Corpus v1.4.0… See the full description on the dataset page: https://huggingface.co/datasets/finaleads/french-corpus-llm-sample.DataShield-Sample-Risk
DataShield
This dataset releases sample-level risk scores for DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment, accepted to the EMNLP Main Conference.
For the method, code, and complete documentation, see the DataShield GitHub repository.
Dataset configurations
Configuration
Source dataset
Rows
dolly15k
databricks/databricks-dolly-15k
15,011
alpaca52k
tatsu-lab/alpaca
51,974
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/killdevil111/DataShield-Sample-Risk.household-samples
WealthSchema Synthetic Household Samples
8 synthetic U.S. households, one per life stage plus one high-net-worth: a small free preview of what a complete, internally consistent household financial profile looks like. Each record covers the people, income, assets, debts, insurance, taxes, goals and a monthly trajectory. No real person is behind any of it.
Built for teams that need realistic households to design, demo or test financial software: planning tools, robo-advisors… See the full description on the dataset page: https://huggingface.co/datasets/wealthschema/household-samples.sample-fusion-intelligence-traces
Sample Fusion Intelligence Traces
Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback.
These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.amazon-reviews-2023-all-beauty-sample
Amazon Reviews 2023 – All_Beauty (Sampled)
This dataset is a sampled subset of the McAuley-Lab/Amazon-Reviews-2023
All_Beauty category, prepared for the YZM2022 Data Mining homework
(Assoc. Prof. Dr. Arzu Kakisim).
Sampling strategy
Source: full All_Beauty reviews (701K) and metadata (112K items).
3-core filtering (each user and item has at least 3 interactions, iterated to convergence).
Cap to the most recent 60 000 interactions, re-applied 3-core.
Metadata restricted… See the full description on the dataset page: https://huggingface.co/datasets/debolut/amazon-reviews-2023-all-beauty-sample.ClimbMix-sample
Unofficial NVIDIA Nemotron-ClimbMix (Subsampled)
This dataset is a curated, subsampled version of OptimalScale/ClimbMix, which itself is a detokenized version of NVIDIA's official pretraining dataset, nvidia/Nemotron-ClimbMix.
It is designed for researchers and developers looking for a smaller, well-shuffled slice of the Nemotron pretraining data for quick experimentation, testing, or ablation studies.
Processing Method
To create this streamlined version, the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ClimbMix-sample.rose_code_samples
rose_code samples (pass@8 rollouts)
vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout
problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass).
Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines.
Qwen3-4B-Thinking-2507/ — teacher model rollouts.
Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8).
Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.reflection-sample-2k
SPP Reflection 2k Sample
A 2,000-row sample (seed 42) of dlab-spp/reflection-10m,
in the identical format, for quick inspection of the data from
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
📦 Full dataset: dlab-spp/reflection-10m (~10M documents).
Each row pairs a pretraining document with a synthetic, value-laden reflection
(first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.GSM-DC-Dataset-Sample
GSM-DC Test Dataset
This dataset contains the test set for GSM-DC (Grade School Math with Distractor Chains), a synthetic math reasoning dataset with controlled complexity.
Dataset Details
Total Problems: 6300
Operation Counts (OP): 16-22 (out-of-distribution test set)
Problem Types: Graph-based mathematical reasoning problems
Noise Levels: Light, Medium, Hard (distractor difficulty)
Dataset Structure
Each problem in all_problems.json contains:
problem_text:… See the full description on the dataset page: https://huggingface.co/datasets/YMinglai/GSM-DC-Dataset-Sample.meetscribe-meeting-samples
MeetScribe Meeting Samples
Synthetic bilingual (EN/FA) enterprise meeting transcripts with labeled action items.
File
Language
Domain
operations_review_en
EN
Production / maintenance
operations_review_en.json
EN
JSON ASR (Whisper format)
safety_board_fa
FA
HSE safety board
procurement_sync_en
EN
Procurement / RFQ
maintenance_planning_fa
FA
Maintenance planning
Usage
python scripts/build_dataset.py
Generates meetings.jsonl with extracted… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/meetscribe-meeting-samples.
