Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ap-mt /flip2-multi-sequence-prompt-ablation-generated-variants-amylase0 likes12k downloads15d agoHugging Face02nganho87098 /sequoia0 likes6.7k downloads14m agoHugging Face03LeeXiangNO1 /DyNativeGaussian_sequence DyNativeGaussian Sequence Demo: Free-Viewpoint Camera Move Dataset Overview DyNativeGaussian_sequence is a curated dynamic scene dataset for research on dynamic scene compression, dynamic novel view synthesis, 4D reconstruction, dynamic Gaussian Splatting, temporal rendering, and video-based scene representation learning. The dataset contains multiple dynamic indoor, outdoor, and performance scenes, including VRU, N3DV, MeetRoom, and Dance-Dunhuang… See the full description on the dataset page: https://huggingface.co/datasets/LeeXiangNO1/DyNativeGaussian_sequence.76 likes6.6k downloads2d agoHugging Face04just-dna-seq /annotators Genomic Variant Annotators Curated genomic variant annotation modules from the DNA-seq project. Overview This dataset contains pre-computed annotation data for genetic variants, organized by module: Module Description Files longevitymap Longevity-associated variants annotations.parquet, studies.parquet, weights.parquet Schema annotations.parquet Variant-level facts linking rsIDs to genes and phenotypes. rsid: dbSNP… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/annotators.image1K<n<10K0 likes4.3k downloads12d agoHugging Face05reeha-parkar /cmu-mosei-comp-seq CMU-MOSEI: Computational Sequences (Unofficial Mirror) This repository provides a mirror of the official computational sequence files from the CMU-MOSEI dataset, which are required for multimodal sentiment and emotion research. The original download links are currently down, so this mirror is provided for the research community. Note: This is an unofficial mirror. All data originates from Carnegie Mellon University and original authors. If you are a dataset creator and want this… See the full description on the dataset page: https://huggingface.co/datasets/reeha-parkar/cmu-mosei-comp-seq.audio5 likes4.2k downloads1y agoHugging Face06ap-mt /flip2-multi-sequence-prompt-ablation-generated-variants-nucb0 likes2.9k downloads14d agoHugging Face07asatheesh /latent-mas-safety-dataset-seq-qwen3-4b LatentMAS Safety Dataset — Phase 0 Latent states, model completions, and safety labels from a Qwen3-4B latent-MAS pipeline (Planner → Critic → Refiner → Judger, inter-agent messages passed as hidden-state vectors) evaluated on prompts from public safety benchmarks. Intended for training a latent safety value model and for probing / interpretability work on multi-agent latent reasoning. What's in it 195,589 rollouts from 14 prompt sources, greedy decode… See the full description on the dataset page: https://huggingface.co/datasets/asatheesh/latent-mas-safety-dataset-seq-qwen3-4b.100K<n<1M0 likes2.6k downloads2mo agoHugging Face08speech-seq2seq /amiGigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription. For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h. For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality.0 likes2.6k downloads4y agoHugging Face09TitouanCh /drug-seq-u2os-novartisI AM NOT AFFILIATED WITH NOVARTIS IN ANY WAY; THIS IS SIMPLY AN UPLOAD OF THEIR DATASET, "NOVARTIS/DRUG-SEQ U2OS MOABOX DATASET." Novartis DRUG-seq U2OS MoABox Dataset This dataset profiles transcriptomic responses of the U-2 OS human osteosarcoma cell line to a broad collection of small molecule perturbations. It contains 49,392 observations spanning 3,742 unique compounds tested at 4 distinct dosages + 0.0, each annotated with their respective mechanisms of action (MoA). Each… See the full description on the dataset page: https://huggingface.co/datasets/TitouanCh/drug-seq-u2os-novartis.text10K<n<100K4 likes2.4k downloads1y agoHugging Face10hlillemark /c4_t5_corrupted_seqlen256 Dataset Card for "c4_t5_corrupted_seqlen256" More Information needed 100M<n<1B0 likes2.3k downloads3y agoHugging Face11klo1 /seq_monkeytext10M<n<100M0 likes1.7k downloads2y agoHugging Face12Shiki42 /PutCab-Sequential-Train50-V4 PutCab Sequential V4 Train50 50 qualified demonstrations from a common 100-scene protocol. Three 320×240 H.264 cameras at 50/3 FPS, 250 Hz physical control, 18143 observations. Construction/packaging Run: E106-R055. The left arm opens the drawer and the right arm grasps, lifts, transfers and releases the object. Sequential runs the two arm programs serially, with half of each order. Concurrent starts both together. CTR keeps frozen normalized Delta choices and records the exact… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/PutCab-Sequential-Train50-V4.videorobotics0 likes1.3k downloads28d agoHugging Face13asatheesh /latent-mas-safety-dataset-seq-qwen3-14b0 likes1.3k downloads1mo agoHugging Face14Shiki42 /PutCab-Sequential-Train100-V4 PutCab Sequential V4 Train100 100 qualified demonstrations from a common 100-scene protocol. Three 320×240 H.264 cameras at 50/3 FPS, 250 Hz physical control, 36795 observations. Construction/packaging Run: E106-R054. The left arm opens the drawer and the right arm grasps, lifts, transfers and releases the object. Sequential runs the two arm programs serially, with half of each order. Concurrent starts both together. CTR keeps frozen normalized Delta choices and records the… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/PutCab-Sequential-Train100-V4.videorobotics0 likes1.1k downloads28d agoHugging Face15t2ance /atlas-25-sequential-tool-runtime-upgrade ATLAS report 25: the sequential tool runtime on verl V1 1. Question and links Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.tabularn<1K0 likes1.1k downloads28d agoHugging Face16Ziyannn /sequential-3d-grounding Sequential 3D Visual Grounding Dataset 5 indoor datasets (ScanNet, HM3D, 3RScan, ARKitScenes, MultiScan) · 10301 scenes · 117885 sequences · 579872 steps Overview Dataset Scenes Train Val Test ScanNet 1513 14501 1700 1808 HM3D 2302 33302 7663 4232 3RScan 1381 13826 4864 1937 ARKitScenes 4834 24349 1394 2653 MultiScan 271 4133 858 665 Total 10301 90111 16479 11295 Annotation Each step has one target (the object to locate)… See the full description on the dataset page: https://huggingface.co/datasets/Ziyannn/sequential-3d-grounding.textother1K<n<10K0 likes1k downloads20d agoHugging Face17Sta8is /cityscapes_seq_video Cityscapes Sequence Video Short video clips built from the Cityscapes sequence data, packaged for training video generation / world models. Contents cityscapes/ ├── train/ │ ├── videos/ 2975 clips (cs_sec_*.mp4) │ ├── metas/ general_prompt.txt — text caption shared by all clips │ └── t5_xxl/ general_prompt.pickle — T5-XXL embedding of that caption └── val/ ├── videos/ 500 clips (cs_sec_*.mp4) ├── metas/… See the full description on the dataset page: https://huggingface.co/datasets/Sta8is/cityscapes_seq_video.videotext-to-video1K<n<10K0 likes961 downloads1mo agoHugging Face18Shiki42 /ctr-scan-object-sequential100-e742-20260923 DEPRECATED / 已废弃(2026-09-24) 数据集采用光栅化渲染,且存在穿模。用户决定废弃该数据集及由其训练的 S015 E745/E747/E749 checkpoint;停止后续评测。历史文件与 revision 保留,仅供溯源,不得作为有效研究结果使用。 Scan Object Sequential100, E742 cohort This LeRobot v3 dataset contains 100 whole Scan Object episodes from E742's 50 new, paired training scenes: one left-first and one right-first episode per scene. The two directions are interleaved by source slot. It uses top and calibrated centered_fovy90 wrist cameras at 25 FPS. No training IdleMask… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-scan-object-sequential100-e742-20260923.robotics0 likes810 downloads16d agoHugging Face19fabricioslv /encode-chip-seq-subset ENCODE Histone ChIP-seq Subset (bigWig signal tracks) Subset of Histone ChIP-seq signal tracks (bigWig format) downloaded from the ENCODE consortium. This is a convenience subset for vectorization / embedding experiments — it is not the full ENCODE release. Contents 46 bigWig files, ~38.7 GB total 3 histone marks: H3K27ac, H3K4me3, H3K27me3 7 experiments (ENCSR accessions): ENCSR349EHZ (10 files) ENCSR491RBV (10 files) ENCSR527FRO (10 files) ENCSR714ZJT (10… See the full description on the dataset page: https://huggingface.co/datasets/fabricioslv/encode-chip-seq-subset.feature-extraction10M<n<100M0 likes692 downloads2mo agoHugging Face20GenerTeam /sequence-recovery Next K-mer Prediction Abouts The Next K-mer Prediction task is a zero-shot evaluation method introduced in the GENERator paper to assess the quality of pretrained models. It involves inputting a sequence segment into the model and having it predict the next K base pairs. The predicted sequence is then compared to the actual sequence to assess accuracy. Sequence: The input sequence has a maximum length of 96k base pairs (bp). You can control the number of input… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/sequence-recovery.texttext-generation10K<n<100K8 likes653 downloads4mo agoHugging Face21Takizz /piperx-water-delivery-0921-75ep-sequential piperx-water-delivery-0921-75ep-sequential 150 episodes from 75 original sources, 171764 frames, 30 FPS. Three camera views and 14-dimensional action/state. Preparation ends at bottle opening +0.6s; synchronization ends at the later tray opening +0.2s. Both independent phases share one first arm and delay ratio. The shared tray motion remains synchronized. observation.arm_active_mask is float32[2], ordered left/right: 0 excludes an optional delay from supervision; 1 retains… See the full description on the dataset page: https://huggingface.co/datasets/Takizz/piperx-water-delivery-0921-75ep-sequential.videon<1K0 likes600 downloads18d agoHugging Face22tomg-group-umd /gemstones_data_order_sequentialGemstones Training Dataset - Sequential version This data is a reprocessed version of the first 1B rows of the Dolma v1.7 dataset (https://huggingface.co/datasets/allenai/dolma). The data is encoded using the Pythia tokenizer: https://huggingface.co/EleutherAI/pythia-160m Disclaimer: this is an approximation of the dataset used to train the Gemstones model suite. Due to the randomized and sharded nature of the distributed training code, the only way to perfectly reproduce the training batches… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/gemstones_data_order_sequential.100M<n<1B0 likes584 downloads1y agoHugging Face23sequelbox /Raiden-DeepSeek-R1Click here to support our open-source dataset and model releases! Raiden-DeepSeek-R1 is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek R1's reasoning skills! This dataset contains: 63k 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1, with all responses generated by deepseek-ai/DeepSeek-R1. Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-DeepSeek-R1.texttext-generation10K<n<100K53 likes582 downloads2y agoHugging Face24Emulated-Inc /api-sequencing-training-pool API sequencing training pool Requests paired with the sequence of API calls that answers them, from three public datasets read at the pinned revisions named below and from a fourth that the builder of this pool generated, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 136292 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file request… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/api-sequencing-training-pool.texttext-generation100K<n<1M0 likes577 downloads28d agoHugging Face25AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes562 downloads24d agoHugging Face26aklein4 /seq2seq-mixed-pretraining-SmolLM2tabular100M<n<1B1 likes537 downloads9mo agoHugging Face27just-dna-seq /alphagenome_avi AlphaGenome AVI scores, re-encoded AlphaGenome's Variant Impact (AVI) scores for 8,812,917,339 SNVs on GRCh38, re-encoded from the 88.5 GB published tabix TSV into ~34 GB of parquet by just-dna-enricher. This is a re-encoding, not a re-analysis. No score is changed, recomputed or filtered. What is in it data/alphagenome_avi-<contig>.parquet chrom, pos (1-based VCF), ref, alt, raw_score_e5 avi_knots.parquet the PHRED reconstruction curve — not… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/alphagenome_avi.tabular1B<n<10B0 likes516 downloads29d agoHugging Face28macwiatrak /bacbench-antibiotic-resistance-protein-sequences Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences) A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces separating different contigs present in the genome. The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.text10K<n<100K0 likes496 downloads1y agoHugging Face29Hiesh /sequence_recovery_centered_200_horizontally RoboMME — SequenceRecoveryHorizontally (Robot HDF5 Demonstrations) Raw robot demonstration data for the SequenceRecoveryHorizontally task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. Each HDF5 file is one recorded episode containing observations (RGB/state), actions, and metadata for imitation learning. Layout train/ 100 episodes val/ 50 episodes test/ 50 episodes Files are named episode_<idx>_seed_<seed>.h5.… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_horizontally.roboticsn<1K0 likes488 downloads3mo agoHugging Face30bloyal /oas-paired-sequence-data Dataset Card for OAS Paired Sequence Data Dataset Summary Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023. textfill-mask1M<n<10M1 likes477 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.