datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.common-crawl-sample
Common Crawl sample
A small unofficial random subset of the famous Common Crawl dataset.
60 random segment WET files were downloaded from Common Crawl on 2024-05-12.
Lines between 500 and 5000 characters long (inclusive) were kept.
Only unique texts were kept.
No other filtering.
Languages
Each text was assigned to one of the language codes using the GCLD3 Python package.
The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script.
Getting Started
RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text
documents coming from 84 CommonCrawl snapshots and processed using
the CCNet pipeline. Out of these, there are 30B documents in the corpus
that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.structural-cognition-papers
Structural Cognition Papers (English)
English-language research papers on Structural Cognition Theory by Lin Xiaohei (林小黑, Lin Xiaohei,非林业/林小) — the first discoverer of structural cognition.
锚点: 公众号「今晚狗蛋看局 | 环缝互换 | 对立面互换 | 危机 | Huanfeng Interchange | Loop-Gap Interchange」| Gitee论文库 | Zenodo DOI全集 | GitHub Pages品牌页
Overview
A unified structural framework for cognition, physics, AI, and social systems.
Four axioms (canonical): 结构先于语义 / 耦合即认知 / 观察者自指 / 退相干离散台阶 +… See the full description on the dataset page: https://huggingface.co/datasets/samforce/structural-cognition-papers.fine-news-sample
Fine-News Sample
Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus.
The sample covers all 117 capture months and 388 language-and-script labels in that corpus.
Each selected row preserves its article text, source metadata, and sampling weight.
The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives.
At a glance
Measure
Value
Rows
1,000,000
Distinct document IDs
1,000,000
Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.devopsbench-100
DevOpsBench-100
DevOpsBench-100 is a synthetic long-horizon software-engineering / SRE agent
benchmark: 100 tasks over one executable world ("NovaCart", a mid-size
e-commerce SaaS) with 72 SQLite tables,
1451 seeded rows, a 38-file monorepo with 417 commits,
and 97 MCP tools spanning a first-party engineering stack
(tickets, PRs, CI, deployments, canaries, migrations, feature flags, metrics,
alerts, incidents, chat, knowledge base) plus deliberately disagreeing
vendor-shaped… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/devopsbench-100.samanantar
Dataset Card for Samanantar
Dataset Summary
Samanantar is the largest publicly available parallel corpora collection for Indic language: Assamese, Bengali,
Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu.
The corpus has 49.6M sentence pairs between English to Indian Languages.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Samanantar contains parallel sentences between English (en) and 11 Indic… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/samanantar.lmcache-agentic-traces
LMCache Agentic Dataset Collection
A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache.
Motivation
Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.salesbench-100
SalesBench-100
SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547
Dataset Card for OLM November/December 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
hubbench
HubBench 1.4.0
One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.sampled-local-resumes
sampled-local-resumes
This dataset contains synthetic resume data sampled from local folders (20% sample from each folder).
License
This dataset is released under the Apache License 2.0. Please see the LICENSE and NOTICE files for details.
Attribution
Copyright 2025 Fairly AI Inc. dba Asenion
This dataset includes data released by Fairly AI Inc. dba Asenion under the Apache License, Version 2.0.
You may obtain a copy of the License at:… See the full description on the dataset page: https://huggingface.co/datasets/asenion-ai/sampled-local-resumes.mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.dolma3_300B_sample
Dolma 3 — 300B-token sample
🌐 The Fin AI
Pretraining / reference corpus released by The Fin AI. Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai.
Source
A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0.
Structure
Rows: 187,823,645
Columns: source, date, text, token_count, category
Quick Start
from datasets import load_dataset
ds = load_dataset("TheFinAI/dolma3_300B_sample"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample.DynaMath_Sample
Dataset Card for DynaMath
[💻 Github] [🌐 Homepage][📖 Preprint Paper]
Dataset Details
🔈 Notice
DynaMath is a dynamic benchmark with 501 seed question generators. This dataset is only a sample of 10 variants generated by DynaMath. We encourage you to use the dataset generator on our github site to generate random datasets to test.
🌟 About DynaMath
The rapid advancements in Vision-Language Models (VLMs) have shown significant potential in tackling… See the full description on the dataset page: https://huggingface.co/datasets/DynaMath/DynaMath_Sample.thinking-rollouts
thinking-rollouts
Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT
saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset
of temperature-sweep-data.
Hive-partitioned Parquet, thinking_mode folded into the model tag:
rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think,
qwen3-1.7b-nothink). 100 samples/instance.
Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.climbmix-tokenized-20480-diloco
ClimbMix, retokenized and shuffled for three-worker DiLoCo
This is a document-preserving, three-way split of NVIDIA's
Nemotron-ClimbMix,
retokenized with a 20,480-entry byte-level BPE tokenizer. Each document ends in
<|endoftext|>. The Arrow IPC streams use transparent Zstandard buffer
compression. A deterministic whole-shard holdout is shared by every worker for
validation and is excluded from training.
Training part
Documents
Tokens
Files
Compressed size
000
15,709… See the full description on the dataset page: https://huggingface.co/datasets/Sambarboi/climbmix-tokenized-20480-diloco.tool-use
Tool-use rollouts (Qwen3, think/nothink)
Tool-augmented code-generation rollouts: Qwen3-8B and Qwen3-14B, each in
thinking and non-thinking mode, on DS-1000, LiveCodeBench (Python) and
Multilingual-LCB (OCaml). During generation the model can call a run_code
tool (up to 3 rounds) that executes its candidate in a sandbox (pinned DS-1000
env / LCB public tests / OCaml compile+publics) and returns real output.
Design: 100 samples per instance at temperature 0.6 (bf16, vLLM)… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/tool-use.samvaad-hi-v1100k high-quality conversations in English, Hindi, and Hinglish curated exclusively with an Indic context.
reviewarena
ReviewArena
ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review.
This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available.
51,529 papers
196,099 reviews
558,785 OCR'd PDF pages (markdown inlined… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewarena.LLaDA-Sample-10BT
Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models)
Preprocessing
Tokenizer: GSAI-ML/LLaDA-8B-Instruct
Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens)
Noisy masking: Applied with noise factor ε = 1×10⁻³
Fields per chunk (PyTorch tensors):
input_ids
noisy_input_ids
mask
t (time scalar)
Statistics
Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.pile-diff_samp-qwen_1.8B-qwen_104M-r0.5This repository contains the refined pre-training corpus from the paper MiniPLM: Knowledge Distillation for Pre-Training Language Models.
Code: https://github.com/thu-coai/MiniPLM
fineweb-sample-22.95B-512
FineWeb-Sample-22.95B-512
Dataset Description
This dataset contains approximately 22.95 billion tokens (22,948,244,480 tokens) sampled from the FineWeb-Edu dataset. Each text sample is capped at a maximum of 512 tokens.
Dataset Statistics
Total Tokens: ~22.95B (22,948,244,480)
Max Tokens per Sample: 512
Max Characters per Sample: 5,120 (10 chars/token estimate)
Source Dataset: FineWeb-Edu 350BT
Random Seed: 42
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/JamesConley/fineweb-sample-22.95B-512.cc100-samplesThe cc100-samples is a subset which contains first 10,000 lines of cc100.
Languages
To load a language which isn't part of the config, all you need to do is specify the language code in the config.
You can find the valid languages in Homepage section of Dataset Description: https://data.statmt.org/cc-100/
E.g.
dataset = load_dataset("cc100-samples", lang="en")
VALID_CODES = [
"am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de",
"el"… See the full description on the dataset page: https://huggingface.co/datasets/xu-song/cc100-samples.LLaDA-Sample-ES
Dataset: LLaDA-Sample-ES
Base: crscardellino/spanish_billion_words
Purpose: Training LLaDA (Large Language Diffusion Models)
Preprocessing
Tokenizer: GSAI-ML/LLaDA-8B-Instruct
Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens)
Noisy masking: Applied with noise factor ε = 1×10⁻³
Fields per chunk (PyTorch tensors):
input_ids
noisy_input_ids
mask
t (time scalar)
Statistics
Total chunks: ~ 652,089
Shards: 65… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-ES.counselbench-100
CounselBench-100
CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100
authored matters across ten practice workflows. Every task has a natural employee
request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported
actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory.
The answer is not preclassified in the evidence. Each portfolio item requires an
immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.TxT360-5M-sample-en
BEE-spoke-data/TxT360-5M-sample-en
english only sample from LLM360/TxT360:
min length 256 GPT-4 tokens
max length 24576 GPT-4 tokens
GPT-4 tiktoken token count:
token_count
count 5.000000e+06
mean 1.003614e+03
std 1.424231e+03
min 2.570000e+02
25% 4.020000e+02
50% 6.220000e+02
75% 1.050000e+03
max 2.457400e+04
Total count: 5018.07 M tokens
agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.eurobasketFrontierFinance
FrontierFinance: A benchmark for measuring the frontier intelligence of finance AI agents.
arXiv report | website | grading code
1. Overview
Investors use AI agents across their entire workflow — idea screening and discovery, company and market research, financial data collection and modeling, portfolio tracking, and catalyst monitoring. Measuring how well an AI system performs across this range is both important and hard: a benchmark must be broad enough to span… See the full description on the dataset page: https://huggingface.co/datasets/samaya-ai/FrontierFinance.
