datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.common-crawl-sample
Common Crawl sample
A small unofficial random subset of the famous Common Crawl dataset.
60 random segment WET files were downloaded from Common Crawl on 2024-05-12.
Lines between 500 and 5000 characters long (inclusive) were kept.
Only unique texts were kept.
No other filtering.
Languages
Each text was assigned to one of the language codes using the GCLD3 Python package.
The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script.
Getting Started
RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text
documents coming from 84 CommonCrawl snapshots and processed using
the CCNet pipeline. Out of these, there are 30B documents in the corpus
that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.fine-news-sample
Fine-News Sample
Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus.
The sample covers all 117 capture months and 388 language-and-script labels in that corpus.
Each selected row preserves its article text, source metadata, and sampling weight.
The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives.
At a glance
Measure
Value
Rows
1,000,000
Distinct document IDs
1,000,000
Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.sampled-local-resumes
sampled-local-resumes
This dataset contains synthetic resume data sampled from local folders (20% sample from each folder).
License
This dataset is released under the Apache License 2.0. Please see the LICENSE and NOTICE files for details.
Attribution
Copyright 2025 Fairly AI Inc. dba Asenion
This dataset includes data released by Fairly AI Inc. dba Asenion under the Apache License, Version 2.0.
You may obtain a copy of the License at:… See the full description on the dataset page: https://huggingface.co/datasets/asenion-ai/sampled-local-resumes.mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.dolma3_300B_sample
Dolma 3 — 300B-token sample
🌐 The Fin AI
Pretraining / reference corpus released by The Fin AI. Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai.
Source
A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0.
Structure
Rows: 187,823,645
Columns: source, date, text, token_count, category
Quick Start
from datasets import load_dataset
ds = load_dataset("TheFinAI/dolma3_300B_sample"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample.DynaMath_Sample
Dataset Card for DynaMath
[💻 Github] [🌐 Homepage][📖 Preprint Paper]
Dataset Details
🔈 Notice
DynaMath is a dynamic benchmark with 501 seed question generators. This dataset is only a sample of 10 variants generated by DynaMath. We encourage you to use the dataset generator on our github site to generate random datasets to test.
🌟 About DynaMath
The rapid advancements in Vision-Language Models (VLMs) have shown significant potential in tackling… See the full description on the dataset page: https://huggingface.co/datasets/DynaMath/DynaMath_Sample.LLaDA-Sample-10BT
Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models)
Preprocessing
Tokenizer: GSAI-ML/LLaDA-8B-Instruct
Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens)
Noisy masking: Applied with noise factor ε = 1×10⁻³
Fields per chunk (PyTorch tensors):
input_ids
noisy_input_ids
mask
t (time scalar)
Statistics
Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.fineweb-sample-22.95B-512
FineWeb-Sample-22.95B-512
Dataset Description
This dataset contains approximately 22.95 billion tokens (22,948,244,480 tokens) sampled from the FineWeb-Edu dataset. Each text sample is capped at a maximum of 512 tokens.
Dataset Statistics
Total Tokens: ~22.95B (22,948,244,480)
Max Tokens per Sample: 512
Max Characters per Sample: 5,120 (10 chars/token estimate)
Source Dataset: FineWeb-Edu 350BT
Random Seed: 42
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/JamesConley/fineweb-sample-22.95B-512.cc100-samplesThe cc100-samples is a subset which contains first 10,000 lines of cc100.
Languages
To load a language which isn't part of the config, all you need to do is specify the language code in the config.
You can find the valid languages in Homepage section of Dataset Description: https://data.statmt.org/cc-100/
E.g.
dataset = load_dataset("cc100-samples", lang="en")
VALID_CODES = [
"am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de",
"el"… See the full description on the dataset page: https://huggingface.co/datasets/xu-song/cc100-samples.LLaDA-Sample-ES
Dataset: LLaDA-Sample-ES
Base: crscardellino/spanish_billion_words
Purpose: Training LLaDA (Large Language Diffusion Models)
Preprocessing
Tokenizer: GSAI-ML/LLaDA-8B-Instruct
Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens)
Noisy masking: Applied with noise factor ε = 1×10⁻³
Fields per chunk (PyTorch tensors):
input_ids
noisy_input_ids
mask
t (time scalar)
Statistics
Total chunks: ~ 652,089
Shards: 65… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-ES.TxT360-5M-sample-en
BEE-spoke-data/TxT360-5M-sample-en
english only sample from LLM360/TxT360:
min length 256 GPT-4 tokens
max length 24576 GPT-4 tokens
GPT-4 tiktoken token count:
token_count
count 5.000000e+06
mean 1.003614e+03
std 1.424231e+03
min 2.570000e+02
25% 4.020000e+02
50% 6.220000e+02
75% 1.050000e+03
max 2.457400e+04
Total count: 5018.07 M tokens
agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tb3-tb4-sample-ml-checkpoint-reshard-recovery
TB3/TB4 Sample Dataset Card
1. Dataset Overview
This repository provides a Harbor terminal-agent benchmark task for the TB3/TB4 Sample stage: ml-checkpoint-reshard-recovery. The task requires an agent to repair an offline distributed-training checkpoint resharding and recovery tool and satisfy an independent, offline, programmatic verifier.
Field
Value
Task ID
ml-checkpoint-reshard-recovery
Primary Domain
ML
Related tags
distributed-systems, storage… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/tb3-tb4-sample-ml-checkpoint-reshard-recovery.evaded-prompt-injection-and-jailbreak-samplesThis dataset originates from our paper 'Bypassing Prompt Injection and Jailbreak Detection in LLM Guardrails'.
The dataset contains a mixture of prompt injections and jailbreak samples modified via character injection and adversarial ML evasion techniques (Techniques can be found within the paper above). For each sample we provide the original unaltered prompt and a modified prompt, the attack_name outlines which attack technique was used to modify the sample.
Acknowledgements… See the full description on the dataset page: https://huggingface.co/datasets/Mindgard/evaded-prompt-injection-and-jailbreak-samples.enterprise-agent-aa-samples
Dataset Card
Dataset Description
Enterprise Agent AA Samples contains three executable enterprise-agent scenarios grounded in frozen public data from UCI, the City of Chicago, and SEC Company Facts. The package combines bilingual task briefs, deterministic stateful environments, normalized tool-use trajectories, source-derived reference outputs, and binary rubric checks.
Task: enterprise tool-use and agent-trajectory evaluation
Languages: English and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/enterprise-agent-aa-samples.dolma3-6t-sample-10000-docs-finance-and-business
HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business
Filename-derived finance_and_business slice of
HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to
revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7.
Extraction rule
The corpus contains every source .jsonl.zst file whose filename contains
the literal segment -finance_and_business-. Source paths and compressed file
contents are preserved byte-for-byte. This is a coarse WebOrganizer
finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.cx_sampled
CulturaX sampled text pools
Use UPLOAD_COMPLETE.json before consuming this release. Its absence means the
upload is incomplete. Pin the completed repository revision for experiments.
Original, unpacked document text from uonlp/CulturaX.
This is a fresh sample, independent of nguyenhuuthuat09/CulturaX_sampled.
No token sequences or training caches are distributed. Each subset has a fixed
validation set and a nested training prefix suitable for smaller token budgets.
Token counts… See the full description on the dataset page: https://huggingface.co/datasets/nht10/cx_sampled.urls-sampled
URLs (hash-sampled)
The same 74,918,894,107 URLs as
ks46/urls, partitioned by
xxh3_64 range into 2,048 chunks of ≈36.6 M rows instead of by SURT key range.
Each chunk is a uniform random sample of the whole corpus, and a URL's chunk
depends on nothing but the URL itself.
Why this exists
The SURT layout groups the web by host: shard 1,000 is a contiguous slice of the
key space, so it holds whole sites and nothing about any other site. That is
what you want for… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-sampled.nemotron-cc-10K-sample-translated
Translated Nemotron-cc-hq samples
This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample
Currently, the following are available, we will add other models and languages:
Model
Languages
Gemma-3-4b-it
["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"]
EuroLLM-9B-Instruct
["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.RedPajama-Data-1T-Sample-Backup
RedPajama Data 1T Sample Backup
This dataset is a backup mirror of togethercomputer/RedPajama-Data-1T-Sample.
It is provided for easier access when the original dataset is unavailable or difficult to download.
Usage
Original:
from datasets import load_dataset
ds = load_dataset(
"togethercomputer/RedPajama-Data-1T-Sample",
split="train",
trust_remote_code=True,
)
Backup:
from datasets import load_dataset
ds = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ll922/RedPajama-Data-1T-Sample-Backup.lldms-associative-memory-samples
LLDMs Associative Memory — Generated Samples
Model-generated text for the paper:
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
Accepted to EMNLP 2026 (Main Conference).
arXiv:2604.26841 · paper · code · checkpoints
29.5 million generated sequences (~3.8B tokens) sampled from the released checkpoints — one
generation run per (model size, training-set fraction). These… See the full description on the dataset page: https://huggingface.co/datasets/lemoncmd/lldms-associative-memory-samples.EDR_Telemetry_SampleThis dataset contains raw Endpoint Detection & Response (EDR) telemetry captured during controlled Deception.Pro malware sandbox operations on an enterprise Active Directory network. Unlike most malware sandboxes — which detonate samples for roughly 30 minutes — our operations run for hours or days per analysis, capturing the full arc of adversary behavior. The data represents a full-fidelity snapshot of system activity recorded while threat actors interacted with a live deception environment… See the full description on the dataset page: https://huggingface.co/datasets/DeceptionPro/EDR_Telemetry_Sample.stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.ultrafineweb-v1.0-sample-100BT
ultrafineweb-v1.0-sample-100BT
A random sample of Ultra-FineWeb v1.0 (English) (files data/ultrafineweb_en/*.parquet at revision 02c85641e3): 60,307,595 of its
1,159,254,991 documents, ~99.9B tokens (estimated at 2.39 UTF-8 bytes per token of an 8k BPE tokenizer), globally shuffled.
split
documents
shards
train
60,248,654
1,023
validation
58,941
1
Columns
text: the document text (the source's content).
score, source: copied unchanged from the… See the full description on the dataset page: https://huggingface.co/datasets/insop/ultrafineweb-v1.0-sample-100BT.JASON-High-Stakes-AI-Evaluation-Samples
J.A.S.O.N. Evaluation Sample Previews V01-V29
Dynamic Response Labs develops specialized data and evaluation resources for high-stakes AI. This public preview introduces the breadth of the J.A.S.O.N. Framework through 29 domain volumes spanning financial stress, operational disruption, coercion and exploitation, cyber incidents, healthcare finance, automated systems, and other consequential contexts.
The collection contains 31 compact preview records. It is designed to help… See the full description on the dataset page: https://huggingface.co/datasets/Dynamicresponselabs/JASON-High-Stakes-AI-Evaluation-Samples.ultrafineweb-l1-hq-sample-100BT
ultrafineweb-l1-hq-sample-100BT
A random sample of Ultra-FineWeb L1 English HQ (files data/ultrafineweb_l1_en_hq/*/*.parquet at revision 02c85641e3): 45,129,641 of its
144,908,921 documents, ~99.8B tokens (estimated at 2.51 UTF-8 bytes per token of an 8k BPE tokenizer), globally shuffled.
split
documents
shards
train
45,085,116
1,023
validation
44,525
1
Columns
text: the document text (the source's content).
meta: copied unchanged from the… See the full description on the dataset page: https://huggingface.co/datasets/insop/ultrafineweb-l1-hq-sample-100BT.
