Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ap-mt /flip2-multi-sequence-prompt-ablation-generated-variants-amylase0 likes8.3k downloads12d agoHugging Face02LeeXiangNO1 /DyNativeGaussian_sequence DyNativeGaussian Sequence Demo: Free-Viewpoint Camera Move Dataset Overview DyNativeGaussian_sequence is a curated dynamic scene dataset for research on dynamic scene compression, dynamic novel view synthesis, 4D reconstruction, dynamic Gaussian Splatting, temporal rendering, and video-based scene representation learning. The dataset contains multiple dynamic indoor, outdoor, and performance scenes, including VRU, N3DV, MeetRoom, and Dance-Dunhuang… See the full description on the dataset page: https://huggingface.co/datasets/LeeXiangNO1/DyNativeGaussian_sequence.76 likes6.8k downloads16d agoHugging Face03ap-mt /flip2-multi-sequence-prompt-ablation-generated-variants-nucb0 likes2.9k downloads12d agoHugging Face04Jules-OC /flowzap-sequence-workflows sequence-workflows A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case. Organization Model Top-level folders are primary Use Cases from the FlowZap Templates dropdown. Second-level folders preserve the original source domain from the FlowZap app index. Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes. Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.tabularn<1K0 likes1.2k downloads6mo agoHugging Face05Hiesh /sequence_recovery_centered_200_horizontally RoboMME — SequenceRecoveryHorizontally (Robot HDF5 Demonstrations) Raw robot demonstration data for the SequenceRecoveryHorizontally task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. Each HDF5 file is one recorded episode containing observations (RGB/state), actions, and metadata for imitation learning. Layout train/ 100 episodes val/ 50 episodes test/ 50 episodes Files are named episode_<idx>_seed_<seed>.h5.… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_horizontally.roboticsn<1K0 likes993 downloads3mo agoHugging Face06Hiesh /sequence_recovery_centered_200_vertically RoboMME — SequenceRecoveryVertically (Robot HDF5 Demonstrations) Raw robot demonstration data for the SequenceRecoveryVertically task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. Each HDF5 file is one recorded episode containing observations (RGB/state), actions, and metadata for imitation learning. Layout train/ 100 episodes val/ 50 episodes test/ 50 episodes Files are named episode_<idx>_seed_<seed>.h5. Seeds… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_vertically.roboticsn<1K0 likes801 downloads3mo agoHugging Face07GenerTeam /sequence-recovery Next K-mer Prediction Abouts The Next K-mer Prediction task is a zero-shot evaluation method introduced in the GENERator paper to assess the quality of pretrained models. It involves inputting a sequence segment into the model and having it predict the next K base pairs. The predicted sequence is then compared to the actual sequence to assess accuracy. Sequence: The input sequence has a maximum length of 96k base pairs (bp). You can control the number of input… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/sequence-recovery.texttext-generation10K<n<100K8 likes667 downloads4mo agoHugging Face08macwiatrak /bacbench-antibiotic-resistance-protein-sequences Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences) A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces separating different contigs present in the genome. The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-protein-sequences.text10K<n<100K0 likes651 downloads1y agoHugging Face09bloyal /oas-paired-sequence-data Dataset Card for OAS Paired Sequence Data Dataset Summary Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023. textfill-mask1M<n<10M1 likes580 downloads3y agoHugging Face10AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes537 downloads22d agoHugging Face11ap-mt /flip2-multi-sequence-prompt-ablation-generated-variants-20 likes385 downloads16d agoHugging Face12AdoCleanCode /sequences_only_correct_V8text1M<n<10M0 likes384 downloads9mo agoHugging Face13AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes374 downloads1mo agoHugging Face14scampion /handball_video_sequencesvideo1K<n<10K0 likes322 downloads1y agoHugging Face15zifeng-ai /genome-sequence-to-function AlphaGenome human compact benchmark This is the materialized human-only 1-Mb benchmark release used by the genome-sequence-to-function track. It contains 3,200 train, 200 validation, and 1,000 test examples; all 5,930 human tracks; and eleven output families. Release ID: release-44d06e9702eba731 Installed size: 92.49 GiB Input context: 1,048,576 bp Target span: 196,608 bp The release keeps the checksummed target objects unchanged and publishes the current public/label/scoring… See the full description on the dataset page: https://huggingface.co/datasets/zifeng-ai/genome-sequence-to-function.other0 likes321 downloads7d agoHugging Face16macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes307 downloads1y agoHugging Face17AL-GR /Origin-Sequence-Data AL-GR/Origin-Sequence-Data: Raw User Behavior Sequences 📜 About the Dataset Each row in this dataset (Origin-Sequence-Data) represents a step in a user's journey, consisting of a sequence of previously interacted items (user_history) and the next item they interacted with (target_item). All item IDs have been anonymized into short, unique strings. This dataset is ideal for: 🧑‍🔬 Researchers who want to design their own data processing or prompting strategies for… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/Origin-Sequence-Data.texttext-generation100K<n<1M0 likes295 downloads1y agoHugging Face18cskokgibbs /yeast-gene-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes283 downloads1y agoHugging Face19mbafca2 /bacbench-antibiotic-resistance-protein-sequences Dataset for antibiotic resistance prediction from whole-bacterial genomes (protein sequences) A dataset of 25,032 bacterial genomes across 39 species with antimicrobial resistance labels. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome, with spaces separating different contigs present in the genome. The antimicrobial resistance labels have been extracted from Antibiotic Susceptibility Test (AST) Browser, accessed 23 Oct, 2024.)… See the full description on the dataset page: https://huggingface.co/datasets/mbafca2/bacbench-antibiotic-resistance-protein-sequences.text10K<n<100K0 likes232 downloads8mo agoHugging Face20macwiatrak /bacbench-essential-genes-protein-sequences Dataset for essential genes prediction in bacterial genomes (Protein sequences) A dataset of 169,408 genes with gene essentiality labels (binary) from 51 bacterial genomes across 37 species. The gene essentiality labels have been extracted from the Database of Essential Genes and the protein sequences have been extracted from GenBank. Each row contains protein sequences present in the genome with an associated essentiality label. We excluded duplicates and genomes with incomplete… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-essential-genes-protein-sequences.textn<1K0 likes222 downloads5mo agoHugging Face21macwiatrak /bacbench-phenotypic-traits-protein-sequences Dataset for phenotypic traits prediction from whole-bacterial genomes (protein sequences) A dataset of 24,462 bacterial genomes across 15,477 species with diverse phenotypic traits as labels. The genome protein sequences have been extracted from GenBank. Each row contains a list of protein sequences present in the bacterial genome, ordered by their location on the chromosome and plasmids. The phenotypic traits have been extracted from a number of sources [1, 2, 3] and include a… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-protein-sequences.text10K<n<100K0 likes222 downloads11mo agoHugging Face22another-phytophile /153-angiosperm-species-32k-sequences-shuffledtabular1M<n<10M0 likes217 downloads5mo agoHugging Face23AntibodyGeneration /sabdab_joint_sequences_uniprottext1K<n<10K6 likes202 downloads3y agoHugging Face24viral-data-safety /sequence_homology_based_v2text10M<n<100M0 likes199 downloads5mo agoHugging Face25Tejaskumar /Emergent-NCA-Sequences-5M ✨ Why this dataset? Emergent NCA Sequences 5M generates complex global behaviors entirely from frozen random Neural Cellular Automata. What makes this approach powerful? Controlled Diversity: Each rollout uses a fresh set of random weights, creating massive diversity in dynamical systems without hand-crafting rules. Stable Semantics: Continuous hidden states are compressed into a global 32-token vocabulary (centroids.pt), guaranteeing… See the full description on the dataset page: https://huggingface.co/datasets/Tejaskumar/Emergent-NCA-Sequences-5M.text-generation1M<n<10M8 likes194 downloads4mo agoHugging Face26Hiesh /robomme_sequencerecoveryvertically RoboMME — SequenceRecoveryVertically (Video QA) Video-QA dataset for the SequenceRecoveryVertically task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.tabularvideo-text-to-textn<1K0 likes189 downloads3mo agoHugging Face27Hiesh /robomme_sequencerecoveryhorizontally RoboMME — SequenceRecoveryHorizontally (Video QA) Video-QA dataset for the SequenceRecoveryHorizontally task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.tabularvideo-text-to-textn<1K0 likes181 downloads3mo agoHugging Face28ronig /pdb_sequences PDB Sequences This dataset contains 780,163 protein sequences from the RCCB Protein Data Bank text100K<n<1M0 likes179 downloads3y agoHugging Face29NoraResearchLab /lithology-sequence-benchmark Lithology Sequence Identification Benchmark Evaluation-only benchmark for reconstructing the lithological sequence of a complete well from raw wireline logs A fixed-well benchmark for testing whether machine-learning and AI systems can infer continuous lithological intervals from conventional petrophysical measurements. Overview The Lithology Sequence Identification Benchmark evaluates models on a practical subsurface interpretation problem: Given the… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark.tabularother1M<n<10M1 likes147 downloads6d agoHugging Face30qualiadev /openarm-restock-sequences-canonical-30fps-subtasks-gripper-vlm restock-sequences-canonical-30fps LeRobot v2.1 dataset: 226 episodes, 306218 frames at 30 fps. Robot: openarm_bimanual Cameras: observation.images.context, observation.images.wrist_left, observation.images.wrist_right State/action dim: 16 Load it with the v2.1 tag, which is the revision the training path pins. tabularrobotics100K<n<1M0 likes146 downloads16d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.