datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-bench-dummy-test-datasettokenizers-test-data
tokenizers-test-data
Test and benchmark fixtures for huggingface/tokenizers,
pulled on demand by the repo Makefiles (make test / make bench / make fixtures
via hf download).
Layout
fixtures/ — multilingual + modality corpora for cross-language encode
benchmarks. Organized, documented, and reproducible: see
fixtures/FIXTURES.md for provenance and
fixtures/fixtures_manifest.json for
exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.Scicode-test-data-h5emit-test-dataset
Dataset Card for EMIT-MSeg Dataset
If you use this dataset, please cite our article:
@misc{herec2026fastmethanedetectionpipeline,
title={A Fast Methane Detection Pipeline on Board Satellites Based on Mag1c-SAS and LinkNet},
author={Jonáš Herec and Vít Růžička and Rado Pitoňák and Jan Sedmidubsky},
year={2026},
eprint={2606.03675},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.03675},
}… See the full description on the dataset page: https://huggingface.co/datasets/onboard-coop/emit-test-dataset.open_lm_test_data_v2scarlet-test-datahindi_audio_dataset_testdata-agent-harbor-test
🎯 Data Agent — Harbor (test)
A held-out benchmark for data-analysis agents: 250 tasks, deliberately balanced across
difficulty and leaning toward the harder end so it actually separates good agents from great ones.
Each task hands your agent a real dataset and a question; it explores, computes, and answers — and
every answer is checked deterministically, with no LLM judge.
Packaged in Harbor format, ready to run.
Where it comes from
Built from the jupyter-agent… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-harbor-test.Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.testDatasetdataset-test-1protein_data_testsplit 1, 2 -> for sequences
split 3, 4 -> for residues
huggingface-testdataworldrenderer-dataset-test
WorldRenderer Dataset Test
Dataset Summary
WorldRenderer Dataset Test is a synthetic multi-scene 3D rendering dataset designed for research in:
Novel View Synthesis (NVS)
Neural Rendering
Geometry-aware Generation
Multi-view Representation Learning
World Models
3D-conditioned Generative Modeling
The dataset contains 27 textured 3D scenes.Each scene is rendered using a predefined monocular camera trajectory consisting of 401 frames.
For every frame, aligned multi-modal… See the full description on the dataset page: https://huggingface.co/datasets/Tengpaz/worldrenderer-dataset-test.OJBench_testdata
🔗 Related Project: OJBench
This dataset is part of the OJBench project — a comprehensive benchmark designed to evaluate large language models on competition-level programming tasks.
📄 Paper: OJBench: A Competition Level Code Benchmark For Large Language Models
💻 Codebase: github.com/He-Ren/OJBench
OJBench focuses on real-world programming contests, featuring 232 curated problems from China’s National Olympiad in Informatics (NOI) and the International Collegiate Programming… See the full description on the dataset page: https://huggingface.co/datasets/He-Ren/OJBench_testdata.isu-challenge-dataset
Dataset Card for ISU Challenge Dataset
Dataset Summary
ISU Challenge Dataset is a multi-modal in-cabin automotive dataset with a controlled synthetic core and paired real-reference scenes.
The synthetic core contains 1,000 synchronized Blender-rendered samples with:
RGB render
depth (EXR and PNG)
instance segmentation
canny edge map
structured scenario labels
An additional 60 paired real-reference scenes occupy sample_00000 through sample_00059. Each paired… See the full description on the dataset page: https://huggingface.co/datasets/ISU-Test/isu-challenge-dataset.AI2_Alphabot_2_test_data_cable
AI2_Alphabot_2_test_data_cable
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 928
Total Frames: 326536
FPS: 30
Dataset Size: 5.77 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_test_data_cable.SolarWM-Data_test-set-v1
SolarWM Standalone Test Set v1
This repository contains the complete, self-contained SolarWM test set without
the training shards. It includes 1,300 clips from 13 source views (100 per
view), packaged as 60 uncompressed WebDataset tar files totaling approximately
77.4 GB.
The test identities are the current accepted SolarWM standalone evaluation
split, excluding MIND. The release contains 757 xhigh and 543 high samples. All selected
samples have non-empty captions and finite… See the full description on the dataset page: https://huggingface.co/datasets/junchaoh-cs/SolarWM-Data_test-set-v1.smart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test.test-datachain-llm-evaltest_datasettest_data
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
🌐 Project Page •
💻 GitHub •
📖 Paper
RPC-Bench is a fine-grained benchmark for research paper comprehension. It is built from review-rebuttal exchanges of high-quality academic papers and supports both text-only and visual evaluation through complementary paper representations.
Data Structure
RPC-Bench is organized into train, dev, and test subsets. Split assignments… See the full description on the dataset page: https://huggingface.co/datasets/kelseye/test_data.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.data_agent_harbor_test
data_agent_harbor_test
250 deterministic data-analysis tasks for agent RL (held-out test benchmark split). Each task gives an agent a Kaggle dataset and a question; the answer is graded deterministically (exact -> numeric tolerance -> list/percent normalization -> symbolic, no LLM judge).
Difficulty tiers: {'hard': 99, 'easy': 85, 'medium': 66}. Environments build from base image savatar101/env-data-agent-train:base.
Format
Harbor task suite: tasks/<id>/… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_harbor_test.testdataprotein_data_test_2autotrain-data-ethnicity-test_v003
AutoTrain Dataset for project: ethnicity-test_v003
Dataset Description
This dataset has been automatically processed by AutoTrain for project ethnicity-test_v003.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<512x512 RGB PIL image>",
"target": 1
},
{
"image": "<512x512 RGB PIL image>",
"target": 3
}]… See the full description on the dataset page: https://huggingface.co/datasets/cledoux42/autotrain-data-ethnicity-test_v003.factur-x-test-data
Factur-X / ZUGFeRD Test Data
Well-formed UN/CEFACT Cross Industry Invoice XML with synthetic seller, buyer, and totals for e-invoicing tests.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
factur-x-small.json / factur-x-small.csv — 100 rows (documents: 25)
factur-x-medium.json / factur-x-medium.csv — 2,000 rows (documents: 250)
factur-x-large.json / factur-x-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/factur-x-test-data.iso20022-test-data
ISO 20022 Test Data
Well-formed pain.001.001.09 Customer Credit Transfer Initiation XML with synthetic debtors, creditors, and amounts.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iso20022-small.json / iso20022-small.csv — 100 rows (documents: 25)
iso20022-medium.json / iso20022-medium.csv — 2,000 rows (documents: 250)
iso20022-large.json / iso20022-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iso20022-test-data.ELLSA_test_data
ELLSA: End-to-end Listen, Look, Speak and Act
The first end-to-end model that unifies vision, speech, text and actionin a streaming full-duplex framework, enabling joint multimodal perception and concurrent generation.
🧪 Highlights
Full-Duplex Multimodal Interaction: unifies listening, looking, speaking, and acting in a single end-to-end architecture, enabling simultaneous… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/ELLSA_test_data.
