datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-processed
RekaDaily-10k (processed)
Short first-person clips cut from the RekaDaily-10k
recordings —
unscripted daily-life video collected through Claru, Reka's
data collection marketplace, recorded by paid collectors in their own homes and
workplaces on head-mounted and handheld phones, across multiple regions.
Every clip carries one dense caption and a multi-question Q&A exchange
written in the second person ("What am I doing in this video?"), so the corpus
drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.ptb-xl-processedProcessed_Interiorverseparliament_hearings_processed
Preprocessed parliament hearings ASR dataset to truecased form.
Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126
dataset_info:
features:
- name: id
dtype: string
- name: audio
dtype:
audio:
sampling_rate: 16000
- name: transcription
sequence: string
splits:
- name: train
num_bytes: 53645064353.18
num_examples: 191455
- name: test
num_bytes: 740331298.0
num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.processed_datasetsBlendedMVS_processeddrh-System-Prompt-processeddeform360_processedaitw-processed-labeled-full
AiTW Processed Full with App Labels
This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset.
Why This Exists
AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.Processed_interiorverse_85multilingual_librispeech_fr_processed
multilingual_librispeech_fr_processed
Dataset Description
Dataset Summary
The data files can be found on the illuin gcloud instance at this adress: unknown_url
This dataset has been processed from Huggingface Hub dataset facebook/multilingual_librispeech and the config french
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual_librispeech_fr_processed.QVHighlights-zipDAPO-Math-17k-Processed
Dataset Card for DAPO-Math-17k-Processed
This is a processed version of BytedTsinghua-SIA/DAPO-Math-17k where we have:
Deduplicated the prompts
Reformatted the prompts and ground truth answers to be compatible with TRL's GRPO trainer
We have also derived pure English and Chinese subsets.
The full dataset processing logic can be found in create_dataset.py.
If you find this dataset useful in your work, please cite the original source with:
@misc{yu2025dapoopensourcellmreinforcement… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/DAPO-Math-17k-Processed.scannet-processed-testHypersim-Processedavqa-processedProcessBench
ProcessBench
This repository contains the dataset of the ProcessBench benchmark proposed by Qwen Team.
You can refer to our GitHub repository for the evaluation code and the prompt templates we use in this work.
If you find this work relevant or helpful to your work, please kindly cite us:
@article{processbench,
title={ProcessBench: Identifying Process Errors in Mathematical Reasoning},
author={
Chujie Zheng and Zhenru Zhang and Beichen Zhang and Runji Lin and Keming Lu and… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/ProcessBench.flyingthings3d_processedprocessed_gpt_dataset_big
Dataset Card for "processed_gpt_dataset_big"
More Information needed
RAGTruth-processed
RAGTruth Dataset
Dataset Description
Dataset Summary
The RAGTruth dataset is designed for evaluating hallucinations in text generation models, particularly in retrieval-augmented generation (RAG) contexts. It contains examples of model outputs along with expert annotations indicating whether the outputs contain hallucinations.
Dataset Structure
Each example contains:
A query/question
Context passages
Model output
Hallucination labels (evident… See the full description on the dataset page: https://huggingface.co/datasets/wandb/RAGTruth-processed.ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.processed_vnhnProcessed-Task-Dataset
Robotic Manipulation Datasets for Four Tasks
[Project Page]
[Paper]
[Code]
[Models]
[Raw GoPro Videos]
This repository contains in-the-wild robotic manipulation datasets collected using UMI, and processed through a SLAM pipeline, as described in the paper "Data Scaling Laws in Imitation Learning for Robotic Manipulation". The datasets cover four tasks:
Pour Water
Arrange Mouse
Fold Towel
Unplug Charger
Dataset Folders:
arrange_mouse and pour_water: Each folder contains… See the full description on the dataset page: https://huggingface.co/datasets/Fanqi-Lin/Processed-Task-Dataset.processed-falcon-dutch-datasetQwen-Terminal-ToolBench-Processed-Tokenized
Qwen Terminal ToolBench Processed Datasets
Qwen-family processed/template-applied and selected tokenized terminal datasets.
Contents
qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text
qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text
qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels
qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.ucf-crime-processed-featuresprocessed_acemath_fullsubgraphrag-processed-embprocessed_fake_job_postings
