datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.alphagenome_avi
AlphaGenome AVI scores, re-encoded
AlphaGenome's Variant Impact (AVI) scores for 8,812,917,339 SNVs on GRCh38, re-encoded from
the 88.5 GB published tabix TSV into ~34 GB of parquet by
just-dna-enricher.
This is a re-encoding, not a re-analysis. No score is changed, recomputed or filtered.
What is in it
data/alphagenome_avi-<contig>.parquet
chrom, pos (1-based VCF), ref, alt, raw_score_e5
avi_knots.parquet
the PHRED reconstruction curve — not… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/alphagenome_avi.seq2seq-mixed-pretraining-SmolLM2l0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts
RL training rollouts
l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091_seqmean
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
ensembl_variations
Ensembl Variations (Parquet Format)
This dataset contains Ensembl human genetic variations converted to Parquet format for fast and efficient VCF annotation.
Usage
With Polars (Recommended)
import polars as pl
# Load variants for chromosome 21
df = pl.scan_parquet("hf://datasets/just-dna-seq/ensembl_variations/data/homo_sapiens-chr21.parquet")
# Filter variants by position
variants = df.filter(
(pl.col("POS") >= 10000000) & (pl.col("POS") <= 20000000)… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/ensembl_variations.tactile-mnist-touch-real-seq-t256-320x240Documentation is available at https://github.com/TimSchneider42/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
ref_seq_bacteria_part_2f-cov-l4096-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts
RL training rollouts
f_cov_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_t1_seqmean_verl091
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
carbon-cpu-enriched-sequences-sampledpiperx-sortletter-20260921-140ep-sequential
piperx-sortletter-20260921-140ep-sequential
140 episodes, 145834 frames, 30 FPS, three 640×480 RGB views. LeRobot v3.0, 14-D action and state.
Selection in output order: original episodes 1–70 (1-based); reverse episodes 71–140 (1-based). CTR retains two relative-timing samples per original episode (280 total); the 140ep name refers to the 140 source demonstrations.
Source: Takizz/piperx-sortletter-0911-0917-140ep-raw, pinned revision f3d743be7a21e8f9d0c5dd23c910ca50c0134ad1.… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-sortletter-20260921-140ep-sequential.ref_seq_vertebrate_non_mammal_part_2ref_seq_vertebrate_non_mammal_part_1dl_alchemy_seq9p6m_context1024ESdB-Embeddings-for-Sequential-data-Benchmark
ESdB: Embeddings for Sequential Data Benchmark
ESdB provides reproducible splits, evaluation shifts, and downstream targets
for benchmarking representations of sequential data.
This repository contains benchmark annotations only. It does not redistribute
the original events or input features. Original datasets must be obtained from
their respective sources and can be reproduced with the preprocessing code in
the ESdB repository.
Structure
Each dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/On-Point-Rnd/ESdB-Embeddings-for-Sequential-data-Benchmark.bacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.gr1_arena_sequential_task_replay
GR1 Arena — Ranch Bottle Into Fridge (new camera pose, ego + wrist, replay)
LeRobot-format teleoperation/replay dataset for the GR1 humanoid performing the
put_item_in_fridge_and_close_door task in Isaac Lab Arena.
Task: Place the ranch dressing bottle on the top shelf of the fridge, and
close the fridge door. (object: ranch_dressing_hope_robolab)
What this dataset is
This is a re-rendered / replayed version of the official NVIDIA Arena
dataset. The source… See the full description on the dataset page: https://huggingface.co/datasets/china-sae-robotics/gr1_arena_sequential_task_replay.robomme_sequencerecoveryhorizontally
RoboMME — SequenceRecoveryHorizontally (Video QA)
Video-QA dataset for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.robomme_sequencerecoveryvertically
RoboMME — SequenceRecoveryVertically (Video QA)
Video-QA dataset for the SequenceRecoveryVertically task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. The agent watches a demonstration video,
remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor
cube before pressing a stop button.
Contents
episodes.parquet — 500 train episodes with per-episode metadata (seeds,
difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.ref_seq_bacteria_part_3yeast-gene-sequence-homology-pretokenized-NTmarinfold-exp11-protein-docs-seq
marinfold-exp11-pdocs-seq
Sequence-only derivative of
eczech/marinfold-exp11-protein-docs.
For every row, the document field has been reduced to just the amino-acid sequence
portion: the <begin_sequence> tag followed by the per-residue three-letter tokens
(e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1>
document-type prefix and everything from <begin_statements> onward (contacts and
distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.pgs-catalog
PGS Catalog — Scoring Files & Cleaned Metadata
Complete mirror of PGS Catalog scoring files converted to
Apache Parquet format, together with cleaned and normalised metadata tables.
Built automatically by the just-prs pipeline.
Last updated: 2026-06-14 16:02 UTC
Release Statistics
Metric
Value
Scoring file parquets
5,337
Unique PGS IDs (metadata)
5,337
Genome build
GRCh38
Total scoring data size
52.6 GB
Release timestamp
2026-06-14T16:02:29Z… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/pgs-catalog.openarm-restock-sequences-canonical-30fps-subtasks-gripper-vlm
restock-sequences-canonical-30fps
LeRobot v2.1 dataset: 226 episodes, 306218 frames at 30 fps.
Robot: openarm_bimanual
Cameras: observation.images.context, observation.images.wrist_left, observation.images.wrist_right
State/action dim: 16
Load it with the v2.1 tag, which is the revision the training path pins.
2d_3d_seq_path_spatial_reasoning
Spatial Reasoning Dataset
A synthetic dataset of Hamiltonian path puzzles with rich chain-of-thought reasoning, designed for training and evaluating spatial reasoning in language models.
Overview
Each sample presents a grid-based puzzle where the solver must find a path visiting every cell exactly once, moving only up/down/left/right (plus above/below for 3D). Puzzles span 2D grids (3x3 to 8x8) and 3D cubes (3x3x3 to 4x4x4), covering solvable, impossible, and multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/eousphoros/2d_3d_seq_path_spatial_reasoning.153-angiosperm-species-32k-sequences-shuffledlithology-sequence-benchmark
Lithology Sequence Identification Benchmark
Evaluation-only benchmark for reconstructing the lithological sequence of a complete well from raw wireline logs
A fixed-well benchmark for testing whether machine-learning and AI systems can infer continuous lithological intervals from conventional petrophysical measurements.
Overview
The Lithology Sequence Identification Benchmark evaluates models on a practical subsurface interpretation problem:
Given the… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark.prs-percentilesless-is-moe-s1-calibration-128-seq8192
Less-is-MoE S1K calibration data — 128 samples, seq_length 8192
This is the fixed calibration artifact used to prune GPT-OSS-120B,
Qwen3.5-122B-A10B, and the Gemma-4-26B-A4B causal language tower. It uses the same 128 source rows as the full-length variant:
yentinglin/s1K-1.1-trl-format revision
58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by
Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows.
For pruning, concatenate messages[].content with one… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.nca-paper-share20-seq_len_2048-657M
nca-paper-share20-seq_len_2048-657M
Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002.
Configuration
param
value
grid
12×12… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-share20-seq_len_2048-657M.
