datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProgramBench-Tests
ProgramBench Generated Tests
This dataset contains the AI-generated behavioral test suites used to evaluate model solutions in ProgramBench.
ProgramBench is a benchmark that evaluates whether language models can rebuild programs from scratch. Given only a compiled binary and its documentation, AI agents must architect and implement a complete codebase that reproduces the original program's behavior. These test suites are used to assess whether a candidate solution is behaviorally… See the full description on the dataset page: https://huggingface.co/datasets/programbench/ProgramBench-Tests.host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.Panda-CVL-test
Panda-CVL Test Split
This dataset is based on the paper onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction. Code is available at GitHub.
Overview
Panda-CVL is a token-level correction dataset and benchmark annotated with the onPanda tool.
Given a question-response pair, the model first judges whether the response is acceptable.
If correction is needed, it must locate the first inappropriate token and replace it… See the full description on the dataset page: https://huggingface.co/datasets/diyer22/Panda-CVL-test.UltraData-Math-TEST
UltraData-Math
🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README
UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models.
🆕 What's New… See the full description on the dataset page: https://huggingface.co/datasets/Alga2025/UltraData-Math-TEST.pile_val_test
The Pile: Validation and Test Splits
This repo contains the validation and test splits of The Pile, an 825 GiB English text dataset designed for training large language models.
Files
File
Split
Size
val.jsonl
Validation
1.4 GB
test.jsonl
Test
1.3 GB
Format
Each line is a JSON object with two fields:
{"text": "The document text...", "meta": {"pile_set_name": "Pile-CC"}}
The meta.pile_set_name field indicates which of the 22 constituent… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pile_val_test.dialogsum-test
Dataset Card for DIALOGSum Corpus
Dataset Description
Links
Homepage: https://aclanthology.org/2021.findings-acl.449
Repository: https://github.com/cylnlp/dialogsum
Paper: https://aclanthology.org/2021.findings-acl.449
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/neil-code/dialogsum-test.hardtests_tests
Dataset Card for HARDTESTS Tests
HARDTESTS Tests is the test suite of HARDTESTS, a competitive programming dataset. The dataset contains multiple .parquet files. Each instance contains structured test case objects for one problem. This dataset is presented in the paper HardTests: Synthesizing High-Quality Test Cases for LLM Coding.
Project Page
Data Summary
The test suite is generated using the HARDTESTSGEN pipeline.
The dataset contains the generated test suites of… See the full description on the dataset page: https://huggingface.co/datasets/sigcp/hardtests_tests.WideSeek-R1-test-data
Testing Dataset
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
android_control_test
Processed Android Control Test Set for InfiGUI-R1 Evaluation
Dataset Description
This repository contains the processed test set derived from the Android Control dataset by Google Research. It has been specifically prepared for evaluating the performance of our model, InfiGUI-R1.
The InfiGUI-R1 model is detailed in our paper:
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
This dataset facilitates standardized testing and… See the full description on the dataset page: https://huggingface.co/datasets/InfiX-ai/android_control_test.encyclopaedia-britannica-lance-test
Encyclopaedia Britannica (1771-1860) - Lance Format
This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading.
Dataset Details
Total Pages: 155,388
Total Volumes: 195
Format: Lance (columnar format with blob storage for images)
Source: National Library of Scotland (NLS)
License: Public Domain (CC0)
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test.FineFineWeb-test
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.otel-test-snippet-jsonl
⚠️ TEST DATASET - DO NOT USE FOR PRODUCTION
This is a small test snippet for internal validation purposes only.
This dataset contains a subset of OpenTelemetry traces from various LLM inference benchmarks. It is intended for testing dataset infrastructure and should NOT be used for research, benchmarking, or production purposes.
Dataset Structure
The dataset contains OpenTelemetry traces organized by:
Benchmark: appworld, tau2_telecom
Agent Framework: openai_solo… See the full description on the dataset page: https://huggingface.co/datasets/lenadan/otel-test-snippet-jsonl.Methods2Test_java_unit_test_code
Dataset Description
Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods.
It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K
Java open source project hosted on GitHub.
The mapping between test case and focal methods are based heuristics rules and Java developer's best practice.
More information could be found here:
methods2test Github repo
Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.Bread-chatbot-dataset-test
Dataset Card for "Bread-chatbot-dataset-test"
More Information needed
aime_2025
AIME 2025 - Unified Test-Time Scaling Format
This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments.
Dataset Description
Source: MathArena/aime_2025
Size: 30 competition-level mathematics problems
Format: Unified TTS format (question, answer, metadata)
Dataset Structure
Fields
question (string): The mathematical problem statement
answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.pi-mono-test
Coding agent session traces DUPLICATED FROM badlogicgames/pi-mono FOR TESTING PURPOSES. ORIGINAL README BELOW
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-mono.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are… See the full description on the dataset page: https://huggingface.co/datasets/JohnBeanerson/pi-mono-test.123_testThe Fewshot Table dataset consists of tables that naturally occur on the web, that are formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. The dataset consists of approximately 413K tables that are extracted from the WDC Web Table Corpora 2015, which is released under the Apache-2.0 license. The WDC Web Table Corpora "contains vast amounts of HTML tables. [...] The Web Data Commons project extracts relational Web tables from the Common Crawl, the largest and most up-to-date Web corpus that is currently available to the public."COMPASS-Policy-Alignment-Testbed-Dataset
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings.
What is COMPASS?
COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.code_x_glue_cc_cloze_testing_all
Dataset Card for "code_x_glue_cc_cloze_testing_all"
Dataset Summary
CodeXGLUE ClozeTesting-all dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-all
Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.
Here we… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_all.x_dataset_test
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/suul999922/x_dataset_test.test
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/boxin-wbx/test.python-unit-test-training-pool
Python unit test training pool
A pool of public data for training a model to write tests for Python code. It is a
straight collection of open datasets, not a new corpus: every row comes from one of the
sources below, at the revision named, and the only rows removed are the ones an overlap
filter flagged against held-out material this pool is kept separate from.
Every row of the normalised layer pairs a program with tests for it. That is the point of
the pool, and it is why the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-unit-test-training-pool.test0611encyclopaedia-britannica-lance-test2
Encyclopaedia Britannica (1771-1860) - Lance Format
This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading.
Dataset Details
Total Pages: 155,388
Total Volumes: 195
Format: Lance (columnar format with blob storage for images)
Source: National Library of Scotland (NLS)
License: Public Domain (CC0)
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test2.code_x_glue_cc_cloze_testing_maxmin
Dataset Card for "code_x_glue_cc_cloze_testing_maxmin"
Dataset Summary
CodeXGLUE ClozeTesting-maxmin dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-maxmin
Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_maxmin.huatuo26M-testdatasets
Dataset Card for huatuo26M-testdatasets
Dataset Summary
We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper.
We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo26M-testdatasets.MM-Mind2Web-tilde_test_snapshot_20dist
MultiModal-Mind2Web~ (MM-Mind2Web~)
rabbit inc.
[Leaderboard & Blogpost to be released]
Configuration: test split, snapshot with seed 42, 20 distractors
Multimodal-Mind2Web is a dataset proposed by Boyuan et al.. It's designed for the development and evaluation of generalist web agents and includes various action trajectories of humans on real websites.
We've simplified the raw dump from both Multimodal-Mind2Web and Mind2Web into sequences of observation-action pairs. We've… See the full description on the dataset page: https://huggingface.co/datasets/rabbit-hmi/MM-Mind2Web-tilde_test_snapshot_20dist.chi-bench
Clinical Healthcare In-Situ Environment
Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark
What is in this dataset
CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/test-alexpouliquen/chi-bench.chem-rlvr-TEST
ChemBench-RLVR: Comprehensive Chemistry Dataset for Reinforcement Learning from Verifiable Rewards
Dataset Description
ChemBench-RLVR is a high-quality, balanced dataset containing 7,001 question-answer pairs across 14 chemistry task types. This dataset is specifically designed for training language models using Reinforcement Learning from Verifiable Rewards (RLVR), where all answers are computationally verifiable using established cheminformatics tools.
Key… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chem-rlvr-TEST.syntra-testing-evals-v4
SyntraTesting Evals v4
Complete benchmark suite for evaluating AI models on advanced reasoning tasks.
Contents
Split
File
Description
prompts
data/splits/prompts.tar.gz (~60KB)
CMT prompts, coherence structures, drift resilience, logic, ethics
benchmarks
data/splits/benchmarks.tar.gz (~36KB)
ARC, CMT, GSM8K benchmark data and utilities
runs
data/splits/runs.tar.gz (~4.4MB)
Sample evaluation runs with graded results
resources
data/splits/resources.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/syntra-testing-evals-v4.
