datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FoldingTShirt_DualArxR5a_Samples
FoldingTShirt_DualArxR5a_Samples
100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2).
Source
Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP.
Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.enterprise-agent-aa-samples
Dataset Card
Dataset Description
Enterprise Agent AA Samples contains three executable enterprise-agent scenarios grounded in frozen public data from UCI, the City of Chicago, and SEC Company Facts. The package combines bilingual task briefs, deterministic stateful environments, normalized tool-use trajectories, source-derived reference outputs, and binary rubric checks.
Task: enterprise tool-use and agent-trajectory evaluation
Languages: English and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/enterprise-agent-aa-samples.multimodal-video-annotation-samples
Video Annotation Samples – SuperviseLab
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories.
Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.JASON-High-Stakes-AI-Evaluation-Samples
J.A.S.O.N. Evaluation Sample Previews V01-V29
Dynamic Response Labs develops specialized data and evaluation resources for high-stakes AI. This public preview introduces the breadth of the J.A.S.O.N. Framework through 29 domain volumes spanning financial stress, operational disruption, coercion and exploitation, cyber incidents, healthcare finance, automated systems, and other consequential contexts.
The collection contains 31 compact preview records. It is designed to help… See the full description on the dataset page: https://huggingface.co/datasets/Dynamicresponselabs/JASON-High-Stakes-AI-Evaluation-Samples.hermes-agent-trace-samples-2026-06-05
Hermes Agent Raw Session Samples
Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers.
Each file in sessions/ is the exact single-session output from:
hermes sessions export sessions/<session_id>.jsonl --session-id <session_id>
No derived tables, flattened rows, SQLite database, or formatted JSON copies are included.
fineweb-edu-100BT-samples-not-in-10BTmultimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.docflow-invoice-samples-fa
DocFlow Invoice Samples — Persian & Bilingual
Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines.
Published by Aria AI Engineering Team.
Dataset Summary
Property
Value
Samples
50 (synthetic, OCR-friendly)
Languages
Persian (FA), English (EN)
Formats
PNG images + JSON annotations
Use case
Invoice OCR benchmarking, AP automation R&D
Synthetic
Yes — no real PII
Fields Annotated
vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.egocentric-samples
Praxis · Egocentric evaluation samples
Real-world hand, object and tool interaction across six camera views. Explore task recordings with MCAP sensor data, MP4 previews, camera intrinsics and extrinsics, and linked metadata.
Contact Tommy · tommy@praxisrobotics.io for evaluation access and tailored data requirements.
At a glance
Included in this release
Coverage
Samples
735
Mapped activity duration
24.16 hours
Camera views
6 per sample
Primary… See the full description on the dataset page: https://huggingface.co/datasets/tommypraxis/egocentric-samples.tripsapien-ai-itinerary-validation-samples
TripSapien public data
CC BY 4.0 sample data for AI itinerary validation: pasted travel plans, expected validation categories, comparison tables, and prompts that show where TripSapien fits in the AI-travel workflow.
TripSapien: https://www.tripsapien.com
Canonical methodology: https://www.tripsapien.com/research/ai-itinerary-validation
Lineage: TripSapien was previously Tripnostic and ValidaTrip, and originally TripPaste.
Why this exists
AI travel planners write… See the full description on the dataset page: https://huggingface.co/datasets/bingwow/tripsapien-ai-itinerary-validation-samples.nemo-stage1-50M-samples
NeMo Stage1 Pretraining Dataset - 50M Samples
This dataset contains 50 million text samples for NeMo model pretraining (Stage 1). The dataset is organized in chunks for efficient loading and processing.
Dataset Details
Total Samples: ~50,000,000
Format: JSONL (JSON Lines)
Structure: Each sample contains {"id": number, "text": "content"}
Chunks: 47 files (chunk_000.jsonl to chunk_046.jsonl)
Samples per chunk: ~1,000,000
Language: English
Task: Text generation pretraining… See the full description on the dataset page: https://huggingface.co/datasets/ssuresh/nemo-stage1-50M-samples.forge-dataset-prep-samples
Forge Dataset Prep — free sample packs
Three downloadable corpora made from public technical documentation. Each pack is chunked JSONL with source metadata, a QA report and a licence notice, so you can inspect the format and trace a record back to its source. Together the three packs hold 1,195 records from 132 source files.
The same packs are described at https://forgedgoods.org/dataset-prep/.
The licence is per pack, not one licence for the repository. The Ollama and llama.cpp… See the full description on the dataset page: https://huggingface.co/datasets/forgedgoods/forge-dataset-prep-samples.Usenet-Corpus-1980-2013-Full-Samples
Usenet Corpus 1980–2013 — Full (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 (cleaned)
dataset — long-form, pre-web Usenet posts. This repo is a free preview; the full,
commercially-licensed corpus (405.8M posts, 102.5B tokens) is at:
Full cleaned dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full
Threaded companion: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full-Samples.synthetic-self-correction-and-thinking-samples
Self Correction and Thinking
A seed library for training language models to reason with self-correction.
Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant.
The structure at a glance
graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.Usenet-Corpus-1980-2013-Threaded-Samples
Usenet Corpus 1980–2013 — Threaded (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded
dataset: Usenet posts reconstructed into conversations via thread_id,
thread_position, and thread_depth. This repo is a free preview; the full,
commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at:
Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded
Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.Longbench_Samples_Specdecpermitguard-ptw-samples
PermitGuard — Synthetic Bilingual Permit-to-Work Samples
Part of the Aria AI oil, gas & petrochemical technical-validation portfolio
(Aria SafeOps → Control of Work / PTW). Companion model:
alirezaaminzadeh/permitguard-risk-classifier
and Space:
alirezaaminzadeh/permitguard-ptw-risk-classifier.
Data honesty
This corpus is 100% synthetic. There are no real permits, incidents, PII, contractors, or named
facilities. Equipment tags such as T-402 / V-101 / P-205B are… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/permitguard-ptw-samples.TACTBench-Samples
TACTBench Demonstration Samples
This repository contains five full-context demonstration examples from
TACTBench. It does not contain the TACT training set or the remaining hidden
TACTBench evaluation set. The samples use the same full-history representation
as the benchmark evaluation and illustrate direct correction, error
explanation, guided revision, clarification checking, affective feedback, and
retry elicitation.
Data
data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.agent-trace-samples
agent-trace-samples
10 example tool-call traces in the agentsnap format. Each row is a single agent run captured as a normalized JSON trace — input, output, tool calls (with hashed results), and a fingerprint.
Useful for:
Testing trace-diffing libraries (we use it in agentsnap's own test suite)
Demoing how to detect silent regressions in agent pipelines (compare a "good" trace with a regressed one)
Onboarding examples for agent observability tooling
Schema
{… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/agent-trace-samples.LoRA-Samples-Intention-Classifier
Dataset Card for LoRA-Samples-Intention-Classifier
Dataset to fine-tune Qwen3-4B-Instruct-2507-LoRA-Intent-Classifier
Dataset Details
Dataset Description
This dataset includes over 10K samples of prompt-intention id pairs for the AI CS agent generator.It is used to fine-tune a small model that powers this agent, reaching a balance of accuracy, efficiency and cost.
Curated by: Li Tuo
Language(s) (NLP): Chinese (primary), English (partial support)
License:… See the full description on the dataset page: https://huggingface.co/datasets/lituokobe/LoRA-Samples-Intention-Classifier.english-classics-parallel-samples
Booklern English classics: parallel samples
Paragraph-aligned opening passages of public-domain English classics with a
translation into Arabic, Spanish, Japanese, Brazilian Portuguese, Russian, Turkish, Chinese, published by Booklern, a
bilingual book reader for learning English through real books. Each book is
read on Booklern with a sentence-by-sentence translation under the English,
read-aloud audio, a dictionary and vocabulary tools; the rows here are the
same opening… See the full description on the dataset page: https://huggingface.co/datasets/ssergaroo/english-classics-parallel-samples.household-samples
WealthSchema Synthetic Household Samples
8 synthetic U.S. households, one per life stage plus one high-net-worth: a small free preview of what a complete, internally consistent household financial profile looks like. Each record covers the people, income, assets, debts, insurance, taxes, goals and a monthly trajectory. No real person is behind any of it.
Built for teams that need realistic households to design, demo or test financial software: planning tools, robo-advisors… See the full description on the dataset page: https://huggingface.co/datasets/wealthschema/household-samples.ipda-golden-samples
IPDA Golden Samples (2AR + 1AR)
Golden samples for fine-tuning debate models on affirmative rebuttal speeches in IPDA format.
Dataset Description
874 high-quality samples for SFT training:
447 2AR (Second Affirmative Rebuttal)
427 1AR (First Affirmative Rebuttal)
Dataset Sources
Source
Count
Description
iter2_group_c
832
High-scoring (>=0.75) samples from GRPO iteration 2
augmented_claude-opus-4.5
20
Augmented debates generated by Claude Opus 4.5… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-golden-samples.one_question_less_samplesCL-bench-samples
CL-bench samples by Mercor
Dataset Description
CL-bench is a benchmark for evaluating language models' context learning abilities.
Resolving tasks in CL-bench requires models to learn from the provided context, ranging from new domain-specific knowledge, rule systems, and complex procedures to laws derived from empirical data, rather than only relying on pre-trained knowledge.
Dataset Structure
Data Fields
Each sample in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/mercor/CL-bench-samples.talentmatch-resume-samples
TalentMatch Resume Samples
Synthetic enterprise resumes and job descriptions with expert HR rankings for benchmark evaluation.
Contents
screenings.jsonl — model vs expert ranks per JD/resume pair
manifest.json — corpus metadata
benchmark_report.json — reproducible metrics
Usage
import json
with open("screenings.jsonl") as f:
for line in f:
print(json.loads(line))
Built by Aria AI.
Pidgin-QandA-data-samples
Pidgin Question-Answer Dataset (Sample)
Sample dataset: Nigerian Pidgin conversational Q&A for dialogue systems and language modeling
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin Question-Answer Dataset (Sample) is a conversational corpus containing 1,462 question-answer pairs entirely in Nigerian Pidgin English. Created by Bytte AI through AI chatbot interactions with human validation, this sample dataset supports dialogue… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-QandA-data-samples.rose_code_samples
rose_code samples (pass@8 rollouts)
vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout
problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass).
Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines.
Qwen3-4B-Thinking-2507/ — teacher model rollouts.
Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8).
Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.ipda-2ar-golden-samples
IPDA Golden Samples (2AR + 1AR)
Golden samples for fine-tuning debate models on affirmative rebuttal speeches in IPDA format.
Dataset Description
422 high-quality samples for SFT training:
260 2AR (Second Affirmative Rebuttal)
162 1AR (First Affirmative Rebuttal)
Dataset Sources
Model
2AR
1AR
Total
Claude Opus 4.5
100
50
150
GPT-5.2
100
50
150
Claude Sonnet
10
10
20
Claude Haiku
9
9
18
Qwen-ft (debate model)
19
19
38
Qwen-base
16
18
34… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-2ar-golden-samples.
