datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.FineFineWeb-test
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.test
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/boxin-wbx/test.chi-bench
Clinical Healthcare In-Situ Environment
Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark
What is in this dataset
CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/test-alexpouliquen/chi-bench.chem-rlvr-TEST
ChemBench-RLVR: Comprehensive Chemistry Dataset for Reinforcement Learning from Verifiable Rewards
Dataset Description
ChemBench-RLVR is a high-quality, balanced dataset containing 7,001 question-answer pairs across 14 chemistry task types. This dataset is specifically designed for training language models using Reinforcement Learning from Verifiable Rewards (RLVR), where all answers are computationally verifiable using established cheminformatics tools.
Key… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chem-rlvr-TEST.agent-trajectories-swe-bench-test-minus-verified
Agent Trajectories: SWE-bench Test \ Verified — Mixed Teachers (gpt-5.2 / gpt-5-mini)
Summary
Full multi-turn agent trajectories collected from the SWE-bench Test minus Verified split
(i.e., SWE-bench Test instances that are not part of SWE-bench Verified).
Intended for SFT of agent models on coding tasks.
Data Collection
Each trajectory was produced by a GT-aware lookahead agent that, at every turn:
Sampled a candidate response from both gpt-5.2 and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/agent-trajectories-swe-bench-test-minus-verified.math_benchmark_test_saturation
LLM Leaderboard Data for Hendrycks MATH Dataset (2022–2024)
This dataset aggregates yearly performance (2022–2024) of large language models (LLMs) on the Hendrycks MATH benchmark. It is specifically compiled to explore performance evolution, benchmark saturation, parameter scaling trends, and evaluation metrics of foundation models solving complex math word problems.
Original source data: Math Word Problem Solving on MATH (Papers with Code)
About Hendrycks' MATH… See the full description on the dataset page: https://huggingface.co/datasets/nlile/math_benchmark_test_saturation.The_OSHA_Test_Project
The OSHA Test Project
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/The_OSHA_Test_Project.test-subset-classic-eda
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 100 rounds.
Every model turn is one row, including the ones that went nowhere.
What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.ClawBench-test
ClawBench
Can AI Agents Complete Everyday Online Tasks?
ClawBench evaluates AI agents on 153 everyday tasks (such as booking flights, ordering groceries, submitting job applications) across 144 live websites. We capture 5 layers of behavioral data (session replay, screenshots, HTTP traffic, agent reasoning traces, and browser actions), collect human ground-truth for every task, and score with an agentic evaluator that provides step-level traceable diagnostics.
Paper… See the full description on the dataset page: https://huggingface.co/datasets/Duke313/ClawBench-test.luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut,
written without ever seeing that continuation. Intended to be spliced into the document
before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the
chunk. Thoughts are strings, not token ids;… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kagent-humanize-test
kagent-humanize-test
2026-09-26 第四版:C18(人设与规划只做输入、范例按场景挑两条、S6 不默认加 emoji)、C19(正式度、话量、口癖频率三项风格参数)、C20(一条气泡等于一个说话动作:S6 只按模型换行分条,不再按逗号切)之后重跑全量。人设仍为小雅。第一版(整改前)结果保留在指标表「kAgent 整改前」列。
200 段中文私聊对话。客户消息取自公开语料原文(与 Aerdax/reply-agent-humanize-test 同一测试集 kapibala_humanize_test),销售消息全部由 kAgent 服务(POST /v1/reply:generate,kAgent commit b026827 加 C18、C19、C20 工作区改动)顺序生成:每一轮请求的 history 是客户原文加此前 kAgent 生成的草稿气泡;kAgent 弃权(abstain)的轮次不产生销售消息,对话里客户消息会连续出现。只保留草稿,不涉及发送。
子集
子集
每行
行数… See the full description on the dataset page: https://huggingface.co/datasets/Aerdax/kagent-humanize-test.reply-agent-humanize-test
reply-agent-humanize-test
200 段中文私聊对话。客户消息取自公开语料原文,销售消息全部由 claude-sonnet-5 按 prompt_zh.md 顺序生成:每一轮生成时模型看到的历史是客户原文加此前生成的回复。
子集
子集
每行
行数
字段
conversations(默认)
一段对话
200
id、messages(role 为 customer 或 sales,content;sales 每条一个气泡)
meta
一段对话的元数据
200
id、source、real_customer、domain、scene、lang、n_messages、n_customer_turns、n_sales_bubbles、n_generated_turns、model、prompt_version、product_info
generated_replies
一轮生成
1116… See the full description on the dataset page: https://huggingface.co/datasets/Aerdax/reply-agent-humanize-test.hermes-testThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
My Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by nex-agi/nex-n2-pro:free.
Sessions: 2
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/hermes-test.kotlin-test-pairs
KotlinTestPairs
9,856 Kotlin source↔unit-test pairs ("focal method" pairs) mined from permissively
licensed public GitHub code.
Applies the methods2test (MSR 2022) methodology to
Kotlin, and is the sibling of
SwiftTestPairs.
⚠️ This dataset contains NO source code
Rows are references + derived metadata: repository, file paths, content MD5s, and
measured properties. The code stays where it has always been, in the upstream corpus.
This avoids redistributing anyone's… See the full description on the dataset page: https://huggingface.co/datasets/Manju46/kotlin-test-pairs.TestiMole
Dataset Card for TestiMole -- A multi-billion tokens Italian text corpus
Testimole is a large linguistic resource for Italian obtained through a massive web scraping effort. As of June 2024, it is one of the largest datasets for the Italian language, if not the largest, publicly available, consisting of almost 100B tokens counted with the Tiktoken cl100k BPE tokenizer. It consists mainly of conversational data (Italian Usenet hierarchies, Italian message boards, Italian… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/TestiMole.test-medicina
Medschool-Test, or "Test di Medicina"
Is your LLM able to pass a National Entrance Exam for the Italian Medical School?
This the GitHub repo for our Hugging Face dataset designed for evaluating Large Language Models (LLMs) on a broad range of questions from the national entrance exams for the Italian medical school (ORIGINAL WEBSITE).
The dataset includes multiple-choice questions from various subjects such as biology, chemistry, physics, mathematics, world… See the full description on the dataset page: https://huggingface.co/datasets/room-b007/test-medicina.8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
Tokenized, tag-wrapped form of JackHsieh/8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.
Each thought is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|/note|>
and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.adaption-floor-test-300
Legal (300-row probe) Requests (Pooled)
Requests from a pooled chat corpus, selected by domain labelling.
Rows
300
Domain
legal (300-row probe)
Format
data.parquet, one row per example
Licence
odc-by
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.
enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-floor-test-300.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.game-of-24
Game of 24 Dataset
Dataset Description
The Game of 24 is a mathematical reasoning puzzle where players must use four numbers and basic arithmetic operations (+, -, *, /) to obtain the result 24. Each number must be used exactly once.
This dataset contains 1,361 unique Game of 24 puzzles ranked by difficulty based on human performance from Amazon Mechanical Turk studies.
Example
Input: 4 5 6 10
Output: (5 * (10 - 4)) - 6 = 24
Step-by-step solution:
10 - 4 = 6… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/game-of-24.sinhala-test-set-50k
Sinhala Test Set - 50K Sentences
A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.4B-general-paragraph-think.stride-32-test.k-8.statml-arxiv.qwen3-ids
4B-general-paragraph-think.stride-32-test.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode. Each record is the model's whole assistant turn: a <think> block, then a note — a long, dense passage of plain prose reasoning about the next 8 tokens after a cut, written from the document prefix alone (the generator never sees the continuation).
The chat template is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-general-paragraph-think.stride-32-test.k-8.statml-arxiv.qwen3-ids.4B-general-paragraph-nothink.stride-32-test.k-8.statml-arxiv.qwen3-ids
4B-general-paragraph-nothink.stride-32-test.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B with thinking turned off. Each thought is a note — a long, dense passage of plain prose reasoning about the next 8 tokens after a cut, written from the document prefix alone (the generator never sees the continuation).
The chat template is domain-free (it never mentions arXiv, papers or… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-general-paragraph-nothink.stride-32-test.k-8.statml-arxiv.qwen3-ids.stratasynth-agent-stress-test
StrataSynth Agent Stress Test
Part of the StrataSynth Synthetic Identity Engineering corpus.
2,068 turns · 100 conversations · 23 columns per turn
High-stakes dialogue explicitly designed to push conversational AI systems. Relationship endings, manager-subordinate conflicts, inheritance disputes. Average relationship tension: 0.69 — the highest of the corpus.
This is where Synthetic Identity Engineering is most visible: under maximum pressure, identities either hold or collapse.… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-agent-stress-test.4B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv
4B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv
Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by
stock Qwen/Qwen3-4B-Instruct-2507 -- no fine-tuning, prompted with the reason-only-nothink.jinja reasoning template.
Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each.
Why this checkpoint: Stock Qwen3-4B-Instruct-2507 prompted with reason-only-nothink.jinja, a byte-identical copy of the reason-only.jinja template used for… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv.swift-test-pairs
SwiftTestPairs
20,801 Swift source↔unit-test pairs ("focal method" pairs) mined from permissively
licensed public GitHub code.
Applies the methods2test (MSR 2022) methodology to
Swift, which had no equivalent dataset.
⚠️ This dataset contains NO source code
Rows are references + derived metadata: repository, file paths, content MD5s, and
measured properties. Reconstruct file contents from the two public upstream datasets with resolve.py
(⚠️ streams ~4 GB from the… See the full description on the dataset page: https://huggingface.co/datasets/Manju46/swift-test-pairs.luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags
luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags
A pre-tokenized, tag-wrapped variant of JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as
<|note|>{thought_text}<|/note|>
and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base token ids (input_ids).
Tag ids and the spliced key are inserted as ids, never re-tokenized.
Delimiter ids: <|note|> = 151669, <|/note|> = 151670. These are the first… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags.test
CS4570 Java Refactoring Benchmark
A benchmark of Java refactoring commits mined from GitHub repositories, split into high-resource and low-resource tiers based on repository popularity. Built for the TU Delft CS4570 ML for Software Engineering course (2026).
Dataset split
Tier
Repositories
Instances
Star range
High-resource
57
618
1000-7298
Low-resource
~60
198
13-90
Total
816
Use the resource_tier field to filter by tier. repo_stars and… See the full description on the dataset page: https://huggingface.co/datasets/Maxros/test.test-traces
