datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.regmix-data-sample
RegMix Data Sample
Dataset Description
The RegMix Data Sample is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task.
Key Features:
Size: Approximately 20GB disk space, 5B tokens
Distribution: Follows the… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data-sample.RoboBrain-X0-Sample-DataThe dataset is currently being uploaded. Please wait a moment.
FoldingTShirt_DualArxR5a_Samples
FoldingTShirt_DualArxR5a_Samples
100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2).
Source
Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP.
Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.tsv_sampleThis folder is the canonical export for the sampled evaluation TSVs.
Files:
HRBench4K.tsv — 300 rows
HRBench8K.tsv — 300 rows
MathVision_MINI.tsv — 300 rows
MathVista_MINI.tsv — 300 rows
MMBench_en_dev.tsv — 300 rows
MME_RealWorld_Lite.tsv — 300 rows
MMMU_val.tsv — 300 rows
MMStar.tsv — 300 rows
MMVet.tsv — 218 rows
POPE.tsv — 300 rows
RealworldQA.tsv — 300 rows
SEED_Bench.tsv — 300 rows
VStarBench.tsv — 191 rows
Notes:
The nine VLMEvalKit-backed TSVs were regenerated from official source… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/tsv_sample.fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.SampledTrajsenterprise-agent-aa-samples
Dataset Card
Dataset Description
Enterprise Agent AA Samples contains three executable enterprise-agent scenarios grounded in frozen public data from UCI, the City of Chicago, and SEC Company Facts. The package combines bilingual task briefs, deterministic stateful environments, normalized tool-use trajectories, source-derived reference outputs, and binary rubric checks.
Task: enterprise tool-use and agent-trajectory evaluation
Languages: English and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/enterprise-agent-aa-samples.dolma3-6t-sample-10000-docs-finance-and-business
HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business
Filename-derived finance_and_business slice of
HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to
revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7.
Extraction rule
The corpus contains every source .jsonl.zst file whose filename contains
the literal segment -finance_and_business-. Source paths and compressed file
contents are preserved byte-for-byte. This is a coarse WebOrganizer
finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.zhongyi-ancient-books-sample-10
未分卷中医药古籍样本数据集
本数据集收录 10 种中医药古籍数字化资料,每种古籍保留一个独立目录,包含 Markdown 识别文本、PDF 原件和分卷表。数据来源于“未分卷中医药古籍 / Part1 / 未分卷ZP”内部整理目录,本次发布为其中 10 个书目的开源样本。
数据不是 JSON-only。仓库中既有结构化清单 JSONL,也有古籍全文 Markdown、PDF 原件和 Excel 分卷表。ModelScope 的 configs 指向 metadata/books_manifest.jsonl,用于数据预览和 SDK 加载;完整古籍文件在 data/books/ 下按书目目录保存。
如需了解更多中医药古籍数字化数据、定制语料整理、OCR 处理、知识库建设或批量授权合作,可发送邮件至 zhouhaoran@shujuyoupu.com。
数据集简介
数据类型:中医药古籍数字化文本、PDF 原件、分卷元数据。
书目数量:10 种。
文件数量:30 个核心文件,另含目录占位文件。… See the full description on the dataset page: https://huggingface.co/datasets/SHPDRG/zhongyi-ancient-books-sample-10.JASON-High-Stakes-AI-Evaluation-Samples
J.A.S.O.N. Evaluation Sample Previews V01-V29
Dynamic Response Labs develops specialized data and evaluation resources for high-stakes AI. This public preview introduces the breadth of the J.A.S.O.N. Framework through 29 domain volumes spanning financial stress, operational disruption, coercion and exploitation, cyber incidents, healthcare finance, automated systems, and other consequential contexts.
The collection contains 31 compact preview records. It is designed to help… See the full description on the dataset page: https://huggingface.co/datasets/Dynamicresponselabs/JASON-High-Stakes-AI-Evaluation-Samples.hermes-agent-trace-samples-2026-06-05
Hermes Agent Raw Session Samples
Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers.
Each file in sessions/ is the exact single-session output from:
hermes sessions export sessions/<session_id>.jsonl --session-id <session_id>
No derived tables, flattened rows, SQLite database, or formatted JSON copies are included.
Arena-DROID-Camera-Sensitivity-Workflow-Sample
Arena DROID Camera Sensitivity Workflow Sample
Dataset Description
Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep.
The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.fineweb-edu-100BT-samples-not-in-10BTmultimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.synthetic-cfo-sample
synthetic cfo public sample: a labelled SAP ECC fiscal year
One complete synthetic company for one fiscal year on SAP ECC table shapes,
generated from accounting rules alone, with fraud planted and recorded in a
ground-truth answer key at the moment it was planted. No real company or
personal data at any stage. There is no original: nothing was sampled, masked
or anonymised.
Engine version 1.11.2. 18 tables, 11,737 rows, 200 labelled fraud records
across 18 distinct schemes.… See the full description on the dataset page: https://huggingface.co/datasets/syntheticcfo/synthetic-cfo-sample.Pandent_Sample
PanDent Sample Release
This repository provides a 300-case sample of PanDent, including 240 cases from the public-source portion of the dataset and 60 de-identified in-house cases.
The sample is released to facilitate inspection of image quality, structured clinical annotations, annotation format, and structure--language correspondence.
The complete PanDent dataset contains 9,524 dental panoramic radiographs, including 9,019 public-source cases and 505 in-house cases.… See the full description on the dataset page: https://huggingface.co/datasets/Desperado1103/Pandent_Sample.docflow-invoice-samples-fa
DocFlow Invoice Samples — Persian & Bilingual
Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines.
Published by Aria AI Engineering Team.
Dataset Summary
Property
Value
Samples
50 (synthetic, OCR-friendly)
Languages
Persian (FA), English (EN)
Formats
PNG images + JSON annotations
Use case
Invoice OCR benchmarking, AP automation R&D
Synthetic
Yes — no real PII
Fields Annotated
vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.egocentric-samples
Praxis · Egocentric evaluation samples
Real-world hand, object and tool interaction across six camera views. Explore task recordings with MCAP sensor data, MP4 previews, camera intrinsics and extrinsics, and linked metadata.
Contact Tommy · tommy@praxisrobotics.io for evaluation access and tailored data requirements.
At a glance
Included in this release
Coverage
Samples
735
Mapped activity duration
24.16 hours
Camera views
6 per sample
Primary… See the full description on the dataset page: https://huggingface.co/datasets/tommypraxis/egocentric-samples.tripsapien-ai-itinerary-validation-samples
TripSapien public data
CC BY 4.0 sample data for AI itinerary validation: pasted travel plans, expected validation categories, comparison tables, and prompts that show where TripSapien fits in the AI-travel workflow.
TripSapien: https://www.tripsapien.com
Canonical methodology: https://www.tripsapien.com/research/ai-itinerary-validation
Lineage: TripSapien was previously Tripnostic and ValidaTrip, and originally TripPaste.
Why this exists
AI travel planners write… See the full description on the dataset page: https://huggingface.co/datasets/bingwow/tripsapien-ai-itinerary-validation-samples.gsma-sample-data
Telecom Benchmark Suite
This repository contains a lightweight benchmarking framework for evaluating Large Language Models (LLMs) on various telecom-related tasks.
It is organized into two main folders:
data/ — contains the datasets in .json format
scripts/ — contains task-specific scripts for prompting and evaluation
Each dataset in data/ has a corresponding Python script in scripts/ that defines:
A system prompt (guidance text for the LLM)
A prompt-building function for… See the full description on the dataset page: https://huggingface.co/datasets/otellm/gsma-sample-data.wmt26-mist-sample
Update Log
22 June 2026 (latest) - we updated our data mix because some BELEBELE samples did not have the context. If you downloaded data before 22 June, please download the new version.
16 June 2026 - first version
Summary
The wmt26-mist-sample is a multilingual mix provided by the WMT26 MIST shared task organizers as a starting point for fine-tuning multilingual LLMs. It contains three types of tasks, to cover same-language and cross-lingual comprehension and… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/wmt26-mist-sample.CLAMP-Sampled-Continuations-and-Demos
CLAMP Sampled Continuations and VLABench Demos
This dataset accompanies HLR/CLAMP, the Constrained Language-Action Model Planner.
It contains 20 closed-loop VLABench case studies: 10 successful and 10 unsuccessful executions. Each case includes the input image and mask, instruction, prompt and entity metadata, CLAMP and ground-truth plans, evaluation output, execution video, and a frame manifest.
Layout
demo_dataset/
├── manifest.json
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/SueMintony/CLAMP-Sampled-Continuations-and-Demos.starcoderdata-samplegsma_sample
GSMA Open-Telco Sample Dataset
Sample data from the GSMA Open-Telco LLM Benchmarks—the first dedicated evaluation framework for assessing LLM performance on telecommunications-specific tasks.
Subsets
Subset
Samples
Task
telemath
100
Telecom-specific mathematical reasoning (signal processing, link budgets, throughput modeling)
teleqna
1,000
Multiple-choice Q&A on telecom standards and domain knowledge
telelogs
100
Root cause analysis for 5G network… See the full description on the dataset page: https://huggingface.co/datasets/emolero/gsma_sample.MMLU_HT_eu_sample
MMLU Human Translated Sample for Basque
A subset of 270 samples manually translated to Basque from the MMLU dataset (Hendrycks et al., 2020). The corresponding 250 English samples are also provided. The MMLU dataset is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/MMLU_HT_eu_sample.nemo-stage1-50M-samples
NeMo Stage1 Pretraining Dataset - 50M Samples
This dataset contains 50 million text samples for NeMo model pretraining (Stage 1). The dataset is organized in chunks for efficient loading and processing.
Dataset Details
Total Samples: ~50,000,000
Format: JSONL (JSON Lines)
Structure: Each sample contains {"id": number, "text": "content"}
Chunks: 47 files (chunk_000.jsonl to chunk_046.jsonl)
Samples per chunk: ~1,000,000
Language: English
Task: Text generation pretraining… See the full description on the dataset page: https://huggingface.co/datasets/ssuresh/nemo-stage1-50M-samples.synthetic-crm-sample
CRM 700 — Free Sample (70 records across 3 tables)
This is a free 70-record sample of the full 700-record commercial dataset. Records are split across three relational tables: customers, products, and orders. Foreign-key relationships are intact — every order references a valid customer and a valid product.
What's in this sample
10 customer records in customers.jsonl
16 product records in products.jsonl
44 order records in orders.jsonl
Referential integrity:… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-crm-sample.
