datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flip2-multi-sequence-prompt-ablation-generated-variants-amylasesequoiaDyNativeGaussian_sequence
DyNativeGaussian Sequence
Demo: Free-Viewpoint Camera Move
Dataset Overview
DyNativeGaussian_sequence is a curated dynamic scene dataset for research on dynamic scene compression, dynamic novel view synthesis, 4D reconstruction, dynamic Gaussian Splatting, temporal rendering, and video-based scene representation learning.
The dataset contains multiple dynamic indoor, outdoor, and performance scenes, including VRU, N3DV, MeetRoom, and Dance-Dunhuang… See the full description on the dataset page: https://huggingface.co/datasets/LeeXiangNO1/DyNativeGaussian_sequence.annotators
Genomic Variant Annotators
Curated genomic variant annotation modules from the DNA-seq project.
Overview
This dataset contains pre-computed annotation data for genetic variants, organized by module:
Module
Description
Files
longevitymap
Longevity-associated variants
annotations.parquet, studies.parquet, weights.parquet
Schema
annotations.parquet
Variant-level facts linking rsIDs to genes and phenotypes.
rsid: dbSNP… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/annotators.cmu-mosei-comp-seq
CMU-MOSEI: Computational Sequences (Unofficial Mirror)
This repository provides a mirror of the official computational sequence files from the CMU-MOSEI dataset, which are required for multimodal sentiment and emotion research. The original download links are currently down, so this mirror is provided for the research community.
Note: This is an unofficial mirror. All data originates from Carnegie Mellon University and original authors. If you are a dataset creator and want this… See the full description on the dataset page: https://huggingface.co/datasets/reeha-parkar/cmu-mosei-comp-seq.flip2-multi-sequence-prompt-ablation-generated-variants-nucblatent-mas-safety-dataset-seq-qwen3-4b
LatentMAS Safety Dataset — Phase 0
Latent states, model completions, and safety labels from a Qwen3-4B
latent-MAS pipeline (Planner → Critic → Refiner → Judger, inter-agent
messages passed as hidden-state vectors) evaluated on prompts from
public safety benchmarks. Intended for training a latent safety value
model and for probing / interpretability work on multi-agent latent
reasoning.
What's in it
195,589 rollouts from 14 prompt sources, greedy decode… See the full description on the dataset page: https://huggingface.co/datasets/asatheesh/latent-mas-safety-dataset-seq-qwen3-4b.amiGigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality
labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised
and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts
and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science,
sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable
for speech recognition training, and to filter out segments with low-quality transcription. For system training,
GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h.
For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage,
and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand,
are re-processed by professional human transcribers to ensure high transcription quality.drug-seq-u2os-novartisI AM NOT AFFILIATED WITH NOVARTIS IN ANY WAY; THIS IS SIMPLY AN UPLOAD OF THEIR DATASET, "NOVARTIS/DRUG-SEQ U2OS MOABOX DATASET."
Novartis DRUG-seq U2OS MoABox Dataset
This dataset profiles transcriptomic responses of the U-2 OS human osteosarcoma cell line to a broad collection of small molecule perturbations. It contains 49,392 observations spanning 3,742 unique compounds tested at 4 distinct dosages + 0.0, each annotated with their respective mechanisms of action (MoA).
Each… See the full description on the dataset page: https://huggingface.co/datasets/TitouanCh/drug-seq-u2os-novartis.c4_t5_corrupted_seqlen256
Dataset Card for "c4_t5_corrupted_seqlen256"
More Information needed
seq_monkeyPutCab-Sequential-Train50-V4
PutCab Sequential V4 Train50
50 qualified demonstrations from a common 100-scene
protocol. Three 320×240 H.264 cameras at 50/3 FPS, 250 Hz physical control,
18143 observations. Construction/packaging Run: E106-R055.
The left arm opens the drawer and the right arm grasps, lifts, transfers and
releases the object. Sequential runs the two arm programs serially, with half
of each order. Concurrent starts both together. CTR keeps frozen normalized
Delta choices and records the exact… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/PutCab-Sequential-Train50-V4.latent-mas-safety-dataset-seq-qwen3-14bPutCab-Sequential-Train100-V4
PutCab Sequential V4 Train100
100 qualified demonstrations from a common 100-scene
protocol. Three 320×240 H.264 cameras at 50/3 FPS, 250 Hz physical control,
36795 observations. Construction/packaging Run: E106-R054.
The left arm opens the drawer and the right arm grasps, lifts, transfers and
releases the object. Sequential runs the two arm programs serially, with half
of each order. Concurrent starts both together. CTR keeps frozen normalized
Delta choices and records the… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/PutCab-Sequential-Train100-V4.atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.sequential-3d-grounding
Sequential 3D Visual Grounding Dataset
5 indoor datasets (ScanNet, HM3D, 3RScan, ARKitScenes, MultiScan) · 10301 scenes · 117885 sequences · 579872 steps
Overview
Dataset
Scenes
Train
Val
Test
ScanNet
1513
14501
1700
1808
HM3D
2302
33302
7663
4232
3RScan
1381
13826
4864
1937
ARKitScenes
4834
24349
1394
2653
MultiScan
271
4133
858
665
Total
10301
90111
16479
11295
Annotation
Each step has one target (the object to locate)… See the full description on the dataset page: https://huggingface.co/datasets/Ziyannn/sequential-3d-grounding.cityscapes_seq_video
Cityscapes Sequence Video
Short video clips built from the Cityscapes
sequence data, packaged for training video generation / world models.
Contents
cityscapes/
├── train/
│ ├── videos/ 2975 clips (cs_sec_*.mp4)
│ ├── metas/ general_prompt.txt — text caption shared by all clips
│ └── t5_xxl/ general_prompt.pickle — T5-XXL embedding of that caption
└── val/
├── videos/ 500 clips (cs_sec_*.mp4)
├── metas/… See the full description on the dataset page: https://huggingface.co/datasets/Sta8is/cityscapes_seq_video.ctr-scan-object-sequential100-e742-20260923
DEPRECATED / 已废弃(2026-09-24)
数据集采用光栅化渲染,且存在穿模。用户决定废弃该数据集及由其训练的 S015 E745/E747/E749 checkpoint;停止后续评测。历史文件与 revision 保留,仅供溯源,不得作为有效研究结果使用。
Scan Object Sequential100, E742 cohort
This LeRobot v3 dataset contains 100 whole Scan Object episodes from E742's 50 new, paired training scenes: one left-first and one right-first episode per scene. The two directions are interleaved by source slot. It uses top and calibrated centered_fovy90 wrist cameras at 25 FPS. No training IdleMask… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-scan-object-sequential100-e742-20260923.encode-chip-seq-subset
ENCODE Histone ChIP-seq Subset (bigWig signal tracks)
Subset of Histone ChIP-seq signal tracks (bigWig format) downloaded from the
ENCODE consortium. This is a convenience subset
for vectorization / embedding experiments — it is not the full ENCODE release.
Contents
46 bigWig files, ~38.7 GB total
3 histone marks: H3K27ac, H3K4me3, H3K27me3
7 experiments (ENCSR accessions):
ENCSR349EHZ (10 files)
ENCSR491RBV (10 files)
ENCSR527FRO (10 files)
ENCSR714ZJT (10… See the full description on the dataset page: https://huggingface.co/datasets/fabricioslv/encode-chip-seq-subset.sequence-recovery
Next K-mer Prediction
Abouts
The Next K-mer Prediction task is a zero-shot evaluation method introduced in the GENERator paper to assess the quality of pretrained models. It involves inputting a sequence segment into the model and having it predict the next K base pairs. The predicted sequence is then compared to the actual sequence to assess accuracy.
Sequence: The input sequence has a maximum length of 96k base pairs (bp). You can control the number of input… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/sequence-recovery.piperx-water-delivery-0921-75ep-sequential
piperx-water-delivery-0921-75ep-sequential
150 episodes from 75 original sources, 171764 frames, 30 FPS. Three camera views and 14-dimensional action/state.
Preparation ends at bottle opening +0.6s; synchronization ends at the later tray opening +0.2s. Both independent phases share one first arm and delay ratio. The shared tray motion remains synchronized.
observation.arm_active_mask is float32[2], ordered left/right: 0 excludes an optional delay from supervision; 1 retains… See the full description on the dataset page: https://huggingface.co/datasets/Takizz/piperx-water-delivery-0921-75ep-sequential.gemstones_data_order_sequentialGemstones Training Dataset - Sequential version
This data is a reprocessed version of the first 1B rows of the Dolma v1.7 dataset (https://huggingface.co/datasets/allenai/dolma).
The data is encoded using the Pythia tokenizer: https://huggingface.co/EleutherAI/pythia-160m
Disclaimer: this is an approximation of the dataset used to train the Gemstones model suite.
Due to the randomized and sharded nature of the distributed training code, the only way to perfectly
reproduce the training batches… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/gemstones_data_order_sequential.Raiden-DeepSeek-R1Click here to support our open-source dataset and model releases!
Raiden-DeepSeek-R1 is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek R1's reasoning skills!
This dataset contains:
63k 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1, with all responses generated by deepseek-ai/DeepSeek-R1.
Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-DeepSeek-R1.api-sequencing-training-pool
API sequencing training pool
Requests paired with the sequence of API calls that answers them, from three public datasets read
at the pinned revisions named below and from a fourth that the builder of this pool generated,
laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 136292 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
request… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/api-sequencing-training-pool.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.seq2seq-mixed-pretraining-SmolLM2alphagenome_avi
AlphaGenome AVI scores, re-encoded
AlphaGenome's Variant Impact (AVI) scores for 8,812,917,339 SNVs on GRCh38, re-encoded from
the 88.5 GB published tabix TSV into ~34 GB of parquet by
just-dna-enricher.
This is a re-encoding, not a re-analysis. No score is changed, recomputed or filtered.
What is in it
data/alphagenome_avi-<contig>.parquet
chrom, pos (1-based VCF), ref, alt, raw_score_e5
avi_knots.parquet
the PHRED reconstruction curve — not… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/alphagenome_avi.bacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.sequence_recovery_centered_200_horizontally
RoboMME — SequenceRecoveryHorizontally (Robot HDF5 Demonstrations)
Raw robot demonstration data for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. Each HDF5 file is one recorded episode
containing observations (RGB/state), actions, and metadata for imitation
learning.
Layout
train/ 100 episodes
val/ 50 episodes
test/ 50 episodes
Files are named episode_<idx>_seed_<seed>.h5.… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_horizontally.oas-paired-sequence-data
Dataset Card for OAS Paired Sequence Data
Dataset Summary
Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023.
