Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01t2ance /atlas-25-sequential-tool-runtime-upgrade ATLAS report 25: the sequential tool runtime on verl V1 1. Question and links Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.tabularn<1K0 likes1.1k downloads28d agoHugging Face02AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes562 downloads25d agoHugging Face03just-dna-seq /alphagenome_avi AlphaGenome AVI scores, re-encoded AlphaGenome's Variant Impact (AVI) scores for 8,812,917,339 SNVs on GRCh38, re-encoded from the 88.5 GB published tabix TSV into ~34 GB of parquet by just-dna-enricher. This is a re-encoding, not a re-analysis. No score is changed, recomputed or filtered. What is in it data/alphagenome_avi-<contig>.parquet chrom, pos (1-based VCF), ref, alt, raw_score_e5 avi_knots.parquet the PHRED reconstruction curve — not… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/alphagenome_avi.tabular1B<n<10B0 likes508 downloads1mo agoHugging Face04aklein4 /seq2seq-mixed-pretraining-SmolLM2tabular100M<n<1B1 likes467 downloads9mo agoHugging Face05hi-todayis-jh /l0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts RL training rollouts l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091_seqmean One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes461 downloads9d agoHugging Face06just-dna-seq /ensembl_variations Ensembl Variations (Parquet Format) This dataset contains Ensembl human genetic variations converted to Parquet format for fast and efficient VCF annotation. Usage With Polars (Recommended) import polars as pl # Load variants for chromosome 21 df = pl.scan_parquet("hf://datasets/just-dna-seq/ensembl_variations/data/homo_sapiens-chr21.parquet") # Filter variants by position variants = df.filter( (pl.col("POS") >= 10000000) & (pl.col("POS") <= 20000000)… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/ensembl_variations.tabulartabular-classification1B<n<10B3 likes455 downloads9mo agoHugging Face07TimSchneider42 /tactile-mnist-touch-real-seq-t256-320x240Documentation is available at https://github.com/TimSchneider42/tactile-mnist/blob/main/doc/datasets.md#touch-datasets. tabularvideo-classificationn<1K0 likes437 downloads1y agoHugging Face08Hack90 /ref_seq_bacteria_part_2tabular100K<n<1M0 likes428 downloads3y agoHugging Face09hi-todayis-jh /f-cov-l4096-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts RL training rollouts f_cov_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_t1_seqmean_verl091 One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes353 downloads8d agoHugging Face10AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes342 downloads2mo agoHugging Face11Shiki42 /piperx-sortletter-20260921-140ep-sequential piperx-sortletter-20260921-140ep-sequential 140 episodes, 145834 frames, 30 FPS, three 640×480 RGB views. LeRobot v3.0, 14-D action and state. Selection in output order: original episodes 1–70 (1-based); reverse episodes 71–140 (1-based). CTR retains two relative-timing samples per original episode (280 total); the 140ep name refers to the 140 source demonstrations. Source: Takizz/piperx-sortletter-0911-0917-140ep-raw, pinned revision f3d743be7a21e8f9d0c5dd23c910ca50c0134ad1.… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-sortletter-20260921-140ep-sequential.tabularrobotics100K<n<1M0 likes331 downloads19d agoHugging Face12Hack90 /ref_seq_vertebrate_non_mammal_part_2tabularn<1K0 likes315 downloads3y agoHugging Face13Hack90 /ref_seq_vertebrate_non_mammal_part_1tabular100K<n<1M0 likes313 downloads3y agoHugging Face14kothasuhas /dl_alchemy_seq9p6m_context1024tabularn<1K0 likes312 downloads1mo agoHugging Face15On-Point-Rnd /ESdB-Embeddings-for-Sequential-data-Benchmark ESdB: Embeddings for Sequential Data Benchmark ESdB provides reproducible splits, evaluation shifts, and downstream targets for benchmarking representations of sequential data. This repository contains benchmark annotations only. It does not redistribute the original events or input features. Original datasets must be obtained from their respective sources and can be reproduced with the preprocessing code in the ESdB repository. Structure Each dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/On-Point-Rnd/ESdB-Embeddings-for-Sequential-data-Benchmark.tabulartabular-classification1M<n<10M3 likes259 downloads3mo agoHugging Face16macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes227 downloads1y agoHugging Face17china-sae-robotics /gr1_arena_sequential_task_replay GR1 Arena — Ranch Bottle Into Fridge (new camera pose, ego + wrist, replay) LeRobot-format teleoperation/replay dataset for the GR1 humanoid performing the put_item_in_fridge_and_close_door task in Isaac Lab Arena. Task: Place the ranch dressing bottle on the top shelf of the fridge, and close the fridge door. (object: ranch_dressing_hope_robolab) What this dataset is This is a re-rendered / replayed version of the official NVIDIA Arena dataset. The source… See the full description on the dataset page: https://huggingface.co/datasets/china-sae-robotics/gr1_arena_sequential_task_replay.tabularrobotics10K<n<100K0 likes225 downloads3mo agoHugging Face18Hiesh /robomme_sequencerecoveryhorizontally RoboMME — SequenceRecoveryHorizontally (Video QA) Video-QA dataset for the SequenceRecoveryHorizontally task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.tabularvideo-text-to-textn<1K0 likes223 downloads3mo agoHugging Face19Hiesh /robomme_sequencerecoveryvertically RoboMME — SequenceRecoveryVertically (Video QA) Video-QA dataset for the SequenceRecoveryVertically task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.tabularvideo-text-to-textn<1K0 likes222 downloads3mo agoHugging Face20Hack90 /ref_seq_bacteria_part_3tabular10K<n<100K0 likes218 downloads3y agoHugging Face21cskokgibbs /yeast-gene-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes210 downloads1y agoHugging Face22eczech /marinfold-exp11-protein-docs-seq marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.tabulartext-generation1M<n<10M0 likes191 downloads4mo agoHugging Face23just-dna-seq /pgs-catalog PGS Catalog — Scoring Files & Cleaned Metadata Complete mirror of PGS Catalog scoring files converted to Apache Parquet format, together with cleaned and normalised metadata tables. Built automatically by the just-prs pipeline. Last updated: 2026-06-14 16:02 UTC Release Statistics Metric Value Scoring file parquets 5,337 Unique PGS IDs (metadata) 5,337 Genome build GRCh38 Total scoring data size 52.6 GB Release timestamp 2026-06-14T16:02:29Z… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/pgs-catalog.tabulartabular-classification10K<n<100K0 likes184 downloads13d agoHugging Face24qualiadev /openarm-restock-sequences-canonical-30fps-subtasks-gripper-vlm restock-sequences-canonical-30fps LeRobot v2.1 dataset: 226 episodes, 306218 frames at 30 fps. Robot: openarm_bimanual Cameras: observation.images.context, observation.images.wrist_left, observation.images.wrist_right State/action dim: 16 Load it with the v2.1 tag, which is the revision the training path pins. tabularrobotics100K<n<1M0 likes182 downloads19d agoHugging Face25eousphoros /2d_3d_seq_path_spatial_reasoning Spatial Reasoning Dataset A synthetic dataset of Hamiltonian path puzzles with rich chain-of-thought reasoning, designed for training and evaluating spatial reasoning in language models. Overview Each sample presents a grid-based puzzle where the solver must find a path visiting every cell exactly once, moving only up/down/left/right (plus above/below for 3D). Puzzles span 2D grids (3x3 to 8x8) and 3D cubes (3x3x3 to 4x4x4), covering solvable, impossible, and multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/eousphoros/2d_3d_seq_path_spatial_reasoning.tabularquestion-answering1K<n<10K0 likes180 downloads8mo agoHugging Face26another-phytophile /153-angiosperm-species-32k-sequences-shuffledtabular1M<n<10M0 likes178 downloads5mo agoHugging Face27NoraResearchLab /lithology-sequence-benchmark Lithology Sequence Identification Benchmark Evaluation-only benchmark for reconstructing the lithological sequence of a complete well from raw wireline logs A fixed-well benchmark for testing whether machine-learning and AI systems can infer continuous lithological intervals from conventional petrophysical measurements. Overview The Lithology Sequence Identification Benchmark evaluates models on a practical subsurface interpretation problem: Given the… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark.tabularother1M<n<10M1 likes160 downloads9d agoHugging Face28just-dna-seq /prs-percentilestabular1M<n<10M0 likes156 downloads1mo agoHugging Face29jayzou3773 /less-is-moe-s1-calibration-128-seq8192 Less-is-MoE S1K calibration data — 128 samples, seq_length 8192 This is the fixed calibration artifact used to prune GPT-OSS-120B, Qwen3.5-122B-A10B, and the Gemma-4-26B-A4B causal language tower. It uses the same 128 source rows as the full-length variant: yentinglin/s1K-1.1-trl-format revision 58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows. For pruning, concatenate messages[].content with one… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.tabulartext-generationn<1K0 likes149 downloads17d agoHugging Face30alexkstern /nca-paper-share20-seq_len_2048-657M nca-paper-share20-seq_len_2048-657M Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002. Configuration param value grid 12×12… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-share20-seq_len_2048-657M.tabularn<1K0 likes146 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.