datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
annotators
Genomic Variant Annotators
Curated genomic variant annotation modules from the DNA-seq project.
Overview
This dataset contains pre-computed annotation data for genetic variants, organized by module:
Module
Description
Files
longevitymap
Longevity-associated variants
annotations.parquet, studies.parquet, weights.parquet
Schema
annotations.parquet
Variant-level facts linking rsIDs to genes and phenotypes.
rsid: dbSNP… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/annotators.BoilingBench-SeqReg
BoilingBench-SeqReg
BoilingBench-SeqReg is the self-contained model-and-test-data package distributed for SeqReg, an open sequence-regression package for experimental pool-boiling heat-flux prediction from hydrophone, AE-hit, and optical-image inputs.
This Hugging Face Dataset preserves the supplied release tree exactly: 79 files totaling approximately 3.12 GiB. It intentionally includes the same four model artifacts hosted at UARK-NED3/SeqReg, so users can obtain the documented… See the full description on the dataset page: https://huggingface.co/datasets/UARK-NED3/BoilingBench-SeqReg.Smoke-Cloud-Segmentation-RACE-ODIN-Data
Smoke Cloud Segmentation Model Training Data
This data is part of the Open Data Integration ODIN (https://nasarace.github.io/race-odin/) project built using the Runtime for Airspace Concept Evaluation (RACE) framework (https://nasarace.github.io/race/)
This is the data for the model here: https://huggingface.co/sequoiaandrade/Smoke-Cloud-Segmentation-RACE-ODIN
The paper for the model is avaialbe here: https://doi.org/10.1016/j.cageo.2025.105960
Copyright (c) 2022, United States… See the full description on the dataset page: https://huggingface.co/datasets/sequoiaandrade/Smoke-Cloud-Segmentation-RACE-ODIN-Data.sim-handover-sequential-200-0820This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda-panda",
"total_episodes": 400,
"total_frames": 74454,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/sim-handover-sequential-200-0820.Seq-DeepFakeseqedge-data
Dataset Card for seqedge-data
Dataset Description
This repository serves as the official public data backend for GalibierHub, an interactive web platform for genomic cohort analytics. The dataset hosts reference bundles, release archives, and sample-level files designed for academic research and downstream bioinformatics pipelines. It currently contains approximately 1.3 GB of open-access data published under the MIT license.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/seqedge-data.rope-diverse-seqThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 25,
"total_frames": 8216,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/rope-diverse-seq.rope-diverse-seq-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5-panda",
"total_episodes": 25,
"total_frames": 8216,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/rope-diverse-seq-v30.1m2r-rope-diverse-seqThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 50,
"total_frames": 16432,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-diverse-seq.Comic-Panel-Sequence-LtoRFor
https://github.com/vsatyamesc/comic-reading-order/tree/main
flock-robotics-vla-training-v2
FLock Robotics VLA Training Dataset v2
Expert demonstrations for the FLock Robotics VLA competition task. All trajectories are successful.
Quick start
from datasets import load_dataset
ds = load_dataset("random-sequence/flock-robotics-vla-training-v2")
print(ds["train"][0])
# Keys: episode_index, step_index, task, difficulty, instruction,
# image (PIL), action [7], proprio [25], reward, done
Dataset statistics
Task
Episodes
Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/random-sequence/flock-robotics-vla-training-v2.cityscapes_sequence_1024by512seqBench
SeqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
Description
SeqBench is a programmatically generated benchmark designed to rigorously evaluate and analyze the sequential reasoning capabilities of language models. Task instances involve pathfinding in 2D grid environments, requiring models to perform multi-step inference over a combination of relevant and distracting textual facts.
The benchmark allows for fine-grained, orthogonal control over… See the full description on the dataset page: https://huggingface.co/datasets/emnlp-submission/seqBench.conslam_seq2_segmentation_pseudoflock-demo-object-detection-sectionsur5-rope-diverse-seqThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 25,
"total_frames": 8216,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/ur5-rope-diverse-seq.panda-rope-diverse-seqThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 25,
"total_frames": 8216,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/panda-rope-diverse-seq.synthetic-cooling-coil-thermal-sequences-v1
Dataset Description
Overview
Name: Synthetic Cooling-Intervention Outcomes v2.
An original reproducible synthetic dataset containing 1,600 thermal observation episodes from 200 equipment units, eight episodes per unit. Each episode has a three-frame thermal raster and a matching visibility raster, plus three alternative future cooling queries, for 4,800 labeled rows and 3,200 PNGs in total. The task is future local temperature-limit exceedance under a specified… See the full description on the dataset page: https://huggingface.co/datasets/darkone01/synthetic-cooling-coil-thermal-sequences-v1.conslam_seq2_segmentation_gtur5-rope-diverse-seq-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5",
"total_episodes": 25,
"total_frames": 8216,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/ur5-rope-diverse-seq-v30.1m2r-rope-diverse-seq-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5-panda",
"total_episodes": 50,
"total_frames": 16432,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/1m2r-rope-diverse-seq-v30.panda-rope-diverse-seq-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 25,
"total_frames": 8216,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/panda-rope-diverse-seq-v30.tbi-sequenceRAW_TIFF_Seqhcs-chess-scoresheets-sequencesseqlab-pooling-benchmark-reportsmessage-decoding-words-and-sequences-target-zoom-in-r1conslam_seq2_19classes_segmentation_gtmessage-decoding-words-and-sequences-target-zoom-inhpatches_sequencesHPatches evaluation dataset
Used by r2d2
