datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flip2-multi-sequence-prompt-ablation-generated-variants-amylaseDyNativeGaussian_sequence
DyNativeGaussian Sequence
Demo: Free-Viewpoint Camera Move
Dataset Overview
DyNativeGaussian_sequence is a curated dynamic scene dataset for research on dynamic scene compression, dynamic novel view synthesis, 4D reconstruction, dynamic Gaussian Splatting, temporal rendering, and video-based scene representation learning.
The dataset contains multiple dynamic indoor, outdoor, and performance scenes, including VRU, N3DV, MeetRoom, and Dance-Dunhuang… See the full description on the dataset page: https://huggingface.co/datasets/LeeXiangNO1/DyNativeGaussian_sequence.flip2-multi-sequence-prompt-ablation-generated-variants-nucbflowzap-sequence-workflows
sequence-workflows
A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case.
Organization Model
Top-level folders are primary Use Cases from the FlowZap Templates dropdown.
Second-level folders preserve the original source domain from the FlowZap app index.
Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes.
Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.sequence_recovery_centered_200_horizontally
RoboMME — SequenceRecoveryHorizontally (Robot HDF5 Demonstrations)
Raw robot demonstration data for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. Each HDF5 file is one recorded episode
containing observations (RGB/state), actions, and metadata for imitation
learning.
Layout
train/ 100 episodes
val/ 50 episodes
test/ 50 episodes
Files are named episode_<idx>_seed_<seed>.h5.… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_horizontally.sequence_recovery_centered_200_vertically
RoboMME — SequenceRecoveryVertically (Robot HDF5 Demonstrations)
Raw robot demonstration data for the SequenceRecoveryVertically task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. Each HDF5 file is one recorded episode
containing observations (RGB/state), actions, and metadata for imitation
learning.
Layout
train/ 100 episodes
val/ 50 episodes
test/ 50 episodes
Files are named episode_<idx>_seed_<seed>.h5. Seeds… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_vertically.sequence-recovery
Next K-mer Prediction
Abouts
The Next K-mer Prediction task is a zero-shot evaluation method introduced in the GENERator paper to assess the quality of pretrained models. It involves inputting a sequence segment into the model and having it predict the next K base pairs. The predicted sequence is then compared to the actual sequence to assess accuracy.
Sequence: The input sequence has a maximum length of 96k base pairs (bp). You can control the number of input… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/sequence-recovery.bacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.oas-paired-sequence-data
Dataset Card for OAS Paired Sequence Data
Dataset Summary
Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023.
carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.flip2-multi-sequence-prompt-ablation-generated-variants-2sequences_only_correct_V8carbon-cpu-enriched-sequences-sampledhandball_video_sequencesgenome-sequence-to-function
AlphaGenome human compact benchmark
This is the materialized human-only 1-Mb benchmark release used by the
genome-sequence-to-function track. It contains 3,200 train, 200 validation,
and 1,000 test examples; all 5,930 human tracks; and eleven output families.
Release ID: release-44d06e9702eba731
Installed size: 92.49 GiB
Input context: 1,048,576 bp
Target span: 196,608 bp
The release keeps the checksummed target objects unchanged and publishes the
current public/label/scoring… See the full description on the dataset page: https://huggingface.co/datasets/zifeng-ai/genome-sequence-to-function.bacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.Origin-Sequence-Data
AL-GR/Origin-Sequence-Data: Raw User Behavior Sequences 📜
About the Dataset
Each row in this dataset (Origin-Sequence-Data) represents a step in a user's journey, consisting of a sequence of previously interacted items (user_history) and the next item they interacted with (target_item). All item IDs have been anonymized into short, unique strings.
This dataset is ideal for:
🧑🔬 Researchers who want to design their own data processing or prompting strategies for… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/Origin-Sequence-Data.yeast-gene-sequence-homology-pretokenized-NTbacbench-antibiotic-resistance-protein-sequences
Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences)
A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces
separating different contigs present in the genome.
The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/mbafca2/bacbench-antibiotic-resistance-protein-sequences.bacbench-essential-genes-protein-sequences
Dataset for essential genes prediction in bacterial genomes (Protein sequences)
A dataset of 169,408 genes with gene essentiality labels (binary) from 51 bacterial genomes across 37 species.
The gene essentiality labels have been extracted from the Database of Essential Genes and the protein sequences have been extracted from GenBank.
Each row contains protein sequences present in the genome with an associated essentiality label. We excluded duplicates and genomes with incomplete… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-essential-genes-protein-sequences.bacbench-phenotypic-traits-protein-sequences
Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences)
A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels.
The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the
bacterial genome, ordered by their location on the chromosome and plasmids.
The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.153-angiosperm-species-32k-sequences-shuffledsabdab_joint_sequences_uniprotsequence_homology_based_v2Emergent-NCA-Sequences-5M
✨ Why this dataset?
Emergent NCA Sequences 5M generates complex global behaviors entirely from frozen random Neural Cellular Automata. What makes this approach powerful?
Controlled Diversity: Each rollout uses a fresh set of random weights, creating massive diversity in dynamical systems without hand-crafting rules.
Stable Semantics: Continuous hidden states are compressed into a global 32-token vocabulary (centroids.pt), guaranteeing… See the full description on the dataset page: https://huggingface.co/datasets/Tejaskumar/Emergent-NCA-Sequences-5M.robomme_sequencerecoveryvertically
RoboMME — SequenceRecoveryVertically (Video QA)
Video-QA dataset for the SequenceRecoveryVertically task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.robomme_sequencerecoveryhorizontally
RoboMME — SequenceRecoveryHorizontally (Video QA)
Video-QA dataset for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.pdb_sequences
PDB Sequences
This dataset contains 780,163 protein sequences from the RCCB Protein Data Bank
lithology-sequence-benchmark
Lithology Sequence Identification Benchmark
Evaluation-only benchmark for reconstructing the lithological sequence of a complete well from raw wireline logs
A fixed-well benchmark for testing whether machine-learning and AI systems can infer continuous lithological intervals from conventional petrophysical measurements.
Overview
The Lithology Sequence Identification Benchmark evaluates models on a practical subsurface interpretation problem:
Given the… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark.openarm-restock-sequences-canonical-30fps-subtasks-gripper-vlm
restock-sequences-canonical-30fps
LeRobot v2.1 dataset: 226 episodes, 306218 frames at 30 fps.
Robot: openarm_bimanual
Cameras: observation.images.context, observation.images.wrist_left, observation.images.wrist_right
State/action dim: 16
Load it with the v2.1 tag, which is the revision the training path pins.
