datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-processed
RekaDaily-10k (processed)
Short first-person clips cut from the RekaDaily-10k
recordings —
unscripted daily-life video collected through Claru, Reka's
data collection marketplace, recorded by paid collectors in their own homes and
workplaces on head-mounted and handheld phones, across multiple regions.
Every clip carries one dense caption and a multi-question Q&A exchange
written in the second person ("What am I doing in this video?"), so the corpus
drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.ptb-xl-processedaitw-processed-labeled-full
AiTW Processed Full with App Labels
This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset.
Why This Exists
AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.processed_vnhnprocessed_fake_job_postingsevalsafe-invoice-processing
Invoice processing
Snapshot: 2026-09-28. 150 cases and 6,874 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.StereoMIS_processeddeepstock-stock-historical-prices-dataset-processedproofwriter_processed_OWAmu-glioma-post-processed
Processed MU-Glioma-Post
Start Here
Use manifests/experiment_index.csv as the main case-level table.
Use manifests/longitudinal_index.csv when ordering patient timepoints for longitudinal work.
Use manifests/validation_summary.json to confirm the processed tree is complete.
Directory Guide
images_native/
Canonical modality symlinks to the source images.
File names use t1, t1c, flair, and t2.
images_reoriented/
Not populated because audit showed all volumes… See the full description on the dataset page: https://huggingface.co/datasets/sbandred/mu-glioma-post-processed.deepstock-stock-historical-prices-dataset-processedthe-stack-v2-processedfannie-mae-processed
Fannie Mae Processed (Parquet)
Dataset respaldado desde Runpod.
Formato principal: Parquet
Ruta de archivos para el visor: loans_locf_parquet/Property_State=/final/.parquet
Nota: la carga excluye temporalmente Property_State=CA.
RoboProcessBench
RoboProcessBench
Dataset Summary
RoboProcessBench is a process-aware benchmark for vision-language robotic manipulation understanding. It evaluates whether VLMs can infer how a manipulation execution unfolds, including phase, contact, motion, bimanual coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions.
This release contains 57,892 QA rows: 48,841 SFT rows and 9,051 evaluation rows across 12 task families and 260 manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ProcessBench-2026/RoboProcessBench.atmosiq-processeddroid-processedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka_panda",
"total_episodes": 60569,
"total_frames": 11936559,
"total_tasks": 34506,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:60569"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/droid-processed.ave-speech-lipemg-processed
AVE Speech — preprocessed for lip–EMG fusion
Code: diddmstjr07/silent-speech-viseme-emg — model definition, training, evaluation, and the analyses behind these numbers.
Related releases: checkpoints · AVE preprocessed · Confusable-100
A derivative of the AVE Speech corpus
(Zhou et al., IEEE THMS 2025), preprocessed into the exact form used to train the
lip–EMG fusion models in the companion work. This is not new recorded data; it is the
original corpus with the preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/diddmstjr/ave-speech-lipemg-processed.robomind_agilex-processedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "agilex_cobot_magic_v2",
"total_episodes": 8638,
"total_frames": 5431759,
"total_tasks": 54,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8638"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/robomind_agilex-processed.processed_test_splitsrh20t_1-processedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "flexiv_Dahuan AG-95",
"total_episodes": 2754,
"total_frames": 959976,
"total_tasks": 124,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:2754"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/rh20t_1-processed.robomind_agilex-trimmed-processedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "agilex_cobot_magic_v2",
"total_episodes": 8638,
"total_frames": 3760151,
"total_tasks": 54,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:8638"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/robomind_agilex-trimmed-processed.qelathrym-hsin-operational-fractal-process-geometry
QELATHRYM–HSIN: Operational Fractal Process Geometry
Task-dependent fractal geometry · Recursive operator trees · Exact task retention · Finite neural tangent representations
Research version 5.0.0 · 1 October 2026 · Self-contained, unreviewed mathematical research draft
QELATHRYM–HSIN proposes a geometry of which quadratic tasks remain accessible through recursive branching under a specified local readout rule. The research connects normalized operator trees, inverse-adjoint… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/qelathrym-hsin-operational-fractal-process-geometry.nyc-tlc-processedmultilingual-amazon-review-sentiment-processedliar2_processed_binaryLOBench-A-share-processedyoutube_processed_full_dataset_finaldocument-processing-benchmark
Document Processing Benchmark
8 public document datasets (receipts, invoices, forms, bank statements,
multi-page docs, contracts) normalized into one parquet schema. Each row
has the document, ground-truth annotations, and per-row token/latency/cost
numbers from real API calls to one or more reference models. You can
read off a target's cost/latency/quality without re-running it.
from datasets import load_dataset
ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.lithium-ion-cell-capacity-grading-process-curves
Lithium-ion Cell Capacity-Grading (FR) Process Curves
Channel-level process curves from the capacity-grading station of a cylindrical
lithium-ion cell line. Every tester channel is sampled natively every 30 seconds for
the whole process, giving the full voltage / current / capacity trajectory of each cell
from the moment it is clamped.
This is the capacity-grading (FR) dataset. Pre-charge is a completely different
process and is published separately; the two are deliberately… See the full description on the dataset page: https://huggingface.co/datasets/michealsmitch/lithium-ion-cell-capacity-grading-process-curves.Multivariate_time_series_data_of_milling_processes_with_varying_tool_wear_and_machine_tools
Flattened version of the "Multivariate time series data of milling processes with varying tool wear and machine tools" dataset in the .parquet file format. Original dataset: https://doi.org/10.17632/zpxs87bjt8.3 and original paper: https://doi.org/10.1016/j.dib.2023.109574
Description
The presented dataset provides labeled, multivariate time series data of milling processes with varying tool wear and for varying machine tools. The width of the flank wear land VB… See the full description on the dataset page: https://huggingface.co/datasets/alpha-by-beta/Multivariate_time_series_data_of_milling_processes_with_varying_tool_wear_and_machine_tools.
