datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-processed
RekaDaily-10k (processed)
Short first-person clips cut from the RekaDaily-10k
recordings —
unscripted daily-life video collected through Claru, Reka's
data collection marketplace, recorded by paid collectors in their own homes and
workplaces on head-mounted and handheld phones, across multiple regions.
Every clip carries one dense caption and a multi-question Q&A exchange
written in the second person ("What am I doing in this video?"), so the corpus
drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.ptb-xl-processedparliament_hearings_processed
Preprocessed parliament hearings ASR dataset to truecased form.
Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126
dataset_info:
features:
- name: id
dtype: string
- name: audio
dtype:
audio:
sampling_rate: 16000
- name: transcription
sequence: string
splits:
- name: train
num_bytes: 53645064353.18
num_examples: 191455
- name: test
num_bytes: 740331298.0
num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.drh-System-Prompt-processedaitw-processed-labeled-full
AiTW Processed Full with App Labels
This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset.
Why This Exists
AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.multilingual_librispeech_fr_processed
multilingual_librispeech_fr_processed
Dataset Description
Dataset Summary
The data files can be found on the illuin gcloud instance at this adress: unknown_url
This dataset has been processed from Huggingface Hub dataset facebook/multilingual_librispeech and the config french
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual_librispeech_fr_processed.DAPO-Math-17k-Processed
Dataset Card for DAPO-Math-17k-Processed
This is a processed version of BytedTsinghua-SIA/DAPO-Math-17k where we have:
Deduplicated the prompts
Reformatted the prompts and ground truth answers to be compatible with TRL's GRPO trainer
We have also derived pure English and Chinese subsets.
The full dataset processing logic can be found in create_dataset.py.
If you find this dataset useful in your work, please cite the original source with:
@misc{yu2025dapoopensourcellmreinforcement… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/DAPO-Math-17k-Processed.scannet-processed-testavqa-processedProcessBench
ProcessBench
This repository contains the dataset of the ProcessBench benchmark proposed by Qwen Team.
You can refer to our GitHub repository for the evaluation code and the prompt templates we use in this work.
If you find this work relevant or helpful to your work, please kindly cite us:
@article{processbench,
title={ProcessBench: Identifying Process Errors in Mathematical Reasoning},
author={
Chujie Zheng and Zhenru Zhang and Beichen Zhang and Runji Lin and Keming Lu and… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/ProcessBench.ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.RAGTruth-processed
RAGTruth Dataset
Dataset Description
Dataset Summary
The RAGTruth dataset is designed for evaluating hallucinations in text generation models, particularly in retrieval-augmented generation (RAG) contexts. It contains examples of model outputs along with expert annotations indicating whether the outputs contain hallucinations.
Dataset Structure
Each example contains:
A query/question
Context passages
Model output
Hallucination labels (evident… See the full description on the dataset page: https://huggingface.co/datasets/wandb/RAGTruth-processed.processed_vnhnprocessed_acemath_fullsft_processed_large_split
sft_processed_large — profile-disjoint split
This is the train / val / test split of Xuhui/sft_processed_large, the
OdysSim midtraining corpus (21.4M interactions across 63 datasets).
Split structure
split
rows
how it's built
train
21.20M
what's left after val + test are carved out
val
28K
per-dataset random sample, in-distribution; for checkpoint selection
test
128K
profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.dna_rendering_processed
DNA-Rendering-Processed Dataset
Project Page | Paper | Code | Model
To enable Diffuman4D model training, we meticulously process the DNA-Rendering dataset by recalibrating camera parameters, optimizing image color correction matrices (CCMs), predicting foreground masks, and estimating human skeletons.
To promote future research in the field of human-centric 3D/4D generation, we have open-sourced our re-annotated labels for the DNA-Rendering dataset in this repo, which includes… See the full description on the dataset page: https://huggingface.co/datasets/krahets/dna_rendering_processed.evalsafe-invoice-processing
Invoice processing
Snapshot: 2026-09-28. 150 cases and 6,874 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.doclaynet_processed
Dataset Card for "doclaynet_processed"
Clean version of DocLayNet ready for finetuning.
finqa-data-processed
FinQA Dataset (Processed)
Dataset Description
Dataset Summary
The FinQA dataset is designed for numerical reasoning over financial data, containing questions that require complex reasoning over tables and text from financial reports.
Dataset Statistics
Total examples: 8281
Training set size: 6624 examples
Test set size: 1657 examples
Dataset Structure
Each example contains:
Required columns:
query: The question to be answered (derived… See the full description on the dataset page: https://huggingface.co/datasets/wandb/finqa-data-processed.coco_val2014_blip2_processed
Dataset Card for "coco_val2014_blip2_processed"
More Information needed
drivaernet_processed
DrivAerNet++ (processed)
This is a processed, downsampled version of DrivAerNet++, not the original dataset.
Fields were converted to a common frame and non-dimensionalized, rows were randomly subsampled and some variables were dropped.
For the original data, see the original paper
The original size around 2.3TB, therefore the volume fields was downsampled 5x, and the surface field is kept as is to reduce the size to 0.78 TB.
Layout
collated/… See the full description on the dataset page: https://huggingface.co/datasets/ayz2/drivaernet_processed.deepstock-stock-historical-prices-dataset-processedVIVOS_CommonVoice_FOSD_Control_processed_dataset
Dataset Card for "VIVOS_CommonVoice_FOSD_Control_processed_dataset"
More Information needed
glotlid_processedopenlegaldata-processed
Dataset Card for openlegaldata.io bulk case data
Dataset Description
This is a edit/cleanup of Bulk Data of openlegaldata.io, which I also brought onto Huggingface here.
The Entire Dataset Is In German
Github Repository: [uniArchive-legalis]](https://github.com/LennardZuendorf/uniArchive-legalis)
Repository: Bulk Data
Edit Summary
I have done some cleaning and splitting of the data and filtered out large parts that were not (easily) usable, cutting… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/openlegaldata-processed.proofwriter_processed_OWAvistr-process-verification-pilot
ViSTR Process-Verification Pilot (14 answer-correct trajectories, multimodal)
Agent trajectories for studying process false positives in multimodal agents:
cases where the answer is correct but the visual reasoning that produced it is
wrong. Ships the raw perception tool outputs so any claim in a trajectory can be
independently re-verified, plus human annotations and an unmodified XSkill
critique of the same trajectories.
Why this exists
Harness / skill… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/vistr-process-verification-pilot.vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
ocr-document-processing-eval
ocr_document_processing_eval
Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks.
Repo: himalaya-ai/ocr-document-processing-eval
Task: document_processing_ocr
Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns.
Optional fine-tuning/eval file: *.sharegpt.json with messages and images.
Core Columns
id: unique sample identifier
image: relative path to the image file
ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.processed_scannet
