Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes562 downloads25d agoHugging Face02AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes342 downloads2mo agoHugging Face03macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes227 downloads1y agoHugging Face04Hiesh /robomme_sequencerecoveryhorizontally RoboMME — SequenceRecoveryHorizontally (Video QA) Video-QA dataset for the SequenceRecoveryHorizontally task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryhorizontally.tabularvideo-text-to-textn<1K0 likes223 downloads3mo agoHugging Face05Hiesh /robomme_sequencerecoveryvertically RoboMME — SequenceRecoveryVertically (Video QA) Video-QA dataset for the SequenceRecoveryVertically task from RoboMME, a ManiSkill/SAPIEN benchmark for memory-augmented robotic manipulation. The agent watches a demonstration video, remembers the arrangement of cubes, and rebuilds it around a pre-placed anchor cube before pressing a stop button. Contents episodes.parquet — 500 train episodes with per-episode metadata (seeds, difficulty, task semantics, language… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/robomme_sequencerecoveryvertically.tabularvideo-text-to-textn<1K0 likes222 downloads3mo agoHugging Face06cskokgibbs /yeast-gene-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes210 downloads1y agoHugging Face07qualiadev /openarm-restock-sequences-canonical-30fps-subtasks-gripper-vlm restock-sequences-canonical-30fps LeRobot v2.1 dataset: 226 episodes, 306218 frames at 30 fps. Robot: openarm_bimanual Cameras: observation.images.context, observation.images.wrist_left, observation.images.wrist_right State/action dim: 16 Load it with the v2.1 tag, which is the revision the training path pins. tabularrobotics100K<n<1M0 likes182 downloads19d agoHugging Face08another-phytophile /153-angiosperm-species-32k-sequences-shuffledtabular1M<n<10M0 likes178 downloads5mo agoHugging Face09NoraResearchLab /lithology-sequence-benchmark Lithology Sequence Identification Benchmark Evaluation-only benchmark for reconstructing the lithological sequence of a complete well from raw wireline logs A fixed-well benchmark for testing whether machine-learning and AI systems can infer continuous lithological intervals from conventional petrophysical measurements. Overview The Lithology Sequence Identification Benchmark evaluates models on a practical subsurface interpretation problem: Given the… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark.tabularother1M<n<10M1 likes160 downloads9d agoHugging Face10Atomi /XES3G5M_interaction_sequencestabular10K<n<100K0 likes119 downloads2y agoHugging Face11neuralbioinfo /PhaStyle-SequenceDB Dataset Card for neuralbioinfo/PhaStyle-SequenceDB phastyle Sequence Database A collection of bacteriophage nucleotide sequences and metadata for training and evaluating phage lifestyle prediction models. Available splits support both strict-holdout and standard-holdout experiments. Dataset Features Name Type Description sequence_id int64 Unique integer identifier for each sequence dataset string Source collection name (see “Splits” below)… See the full description on the dataset page: https://huggingface.co/datasets/neuralbioinfo/PhaStyle-SequenceDB.tabular1K<n<10K0 likes108 downloads1y agoHugging Face12random-sequence /flock-demo-critical-infra-sectionstabular10K<n<100K0 likes75 downloads8mo agoHugging Face13Jules-OC /flowzap-sequence-workflows sequence-workflows A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case. Organization Model Top-level folders are primary Use Cases from the FlowZap Templates dropdown. Second-level folders preserve the original source domain from the FlowZap app index. Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes. Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.tabularn<1K0 likes73 downloads7mo agoHugging Face14random-sequence /flock-video-inconsistency FLock Video Inconsistency Short procedurally generated videos (default 320x240 at 15 fps, 6 to 10 s), each a continuous shot into which 0 to 4 known inconsistencies were injected, together with legitimate, unlabelled decoy events that look like edits but are not. Every clip comes with frame-exact labels. The data trains detectors for the FLock video_inconsistency task: given a video, list every inconsistency with its type, time span, confidence and (for two types) a bounding… See the full description on the dataset page: https://huggingface.co/datasets/random-sequence/flock-video-inconsistency.tabularvideo-classification1K<n<10K0 likes58 downloads10d agoHugging Face15another-phytophile /153-angiosperm-species-8192bp-sequencestabular1M<n<10M0 likes52 downloads6mo agoHugging Face16Aditya02 /Charades-Action-Sequence-Sampletabular1K<n<10K0 likes47 downloads2y agoHugging Face17CodeIsAbstract /sanskrit-morpho-sequences Sanskrit Morphological Sequence Corpus (Vidyut-Verified) A large-scale, Pāṇinian-verified morphological sequence dataset for classical and Vedic Sanskrit. Every token is annotated with its lemma, generative root (aupadeśika), part-of-speech, case, number, person, voice, and gender — all in the SLP1 transliteration, and all aligned at the sentence level for sequence-tagging / seq2seq training. 710,785 sentences (after deduplication) 5,511,664 tokens 14 columns (10 linguistic + 4… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-morpho-sequences.tabulartoken-classification100K<n<1M0 likes47 downloads3mo agoHugging Face18junha1125 /openvid-frame-sequences-1M OpenVid Frame Sequences — 1M adjacent frame pairs Short, single-shot frame sequences cut from OpenVid-1M, built to train and evaluate models on what changes between two frames half a second apart. One sample = 10 consecutive frames, 0.5 s apart (a 4.5 s span) → 9 adjacent frame pairs. [f00] --0.5s--> [f01] --0.5s--> [f02] ... [f09] ^ the thing you describe / predict Sequences 116,596 Frames per sequence 10 (0.5 s apart, t = 0.0 … 4.5 s) Adjacent frame… See the full description on the dataset page: https://huggingface.co/datasets/junha1125/openvid-frame-sequences-1M.tabularimage-to-text100K<n<1M0 likes46 downloads2mo agoHugging Face19XingweiT /IntrEx-sequence IntrEx: A Dataset for Modeling Engagement in Educational Conversations (sequence-level) 【 📦 GitHub repo | 🤗 Paper 】 TL;DR IntrEx is the first large-scale dataset annotated for interestingness and expected interestingness in teacher-student interactions. Data Fields Column Description project_id ID for specifying a unit of annotation work where a batch of participants annotate a set of conversations page_id The annotation page number inside… See the full description on the dataset page: https://huggingface.co/datasets/XingweiT/IntrEx-sequence.tabulartext-classification1K<n<10K1 likes44 downloads1y agoHugging Face20willdaspit /afdb_50_sequence_clustered_reprs_GOGO functional annotations are semicolon-separated in go_ids. The "GO:" prefix is stripped. Only ids appearing at least 10000 times in the training set are kept - there are 425 such ids. Parents are automatically populated (e.g. iron binding -> metal binding). Obsolete ids are replaced with current where possible, or removed if not. tabular10M<n<100M0 likes39 downloads8mo agoHugging Face21willdaspit /afdb_50_sequence_clusteredAnnotated sequences from the AFDB50 (sequence-based, not structure-based) clustering at https://afdb-cluster.steineggerlab.workers.dev/ Two rows with the same RepId are part of the same cluster. Similarly, two rows with the same cluster_id are part of the same cluster. cluster_id is in order from smallest to largest cluster - that is, members of the smallest cluster have cluster_id=0. Singletons are included. All plddts are included. All are annotated with the number of cluster members, plddt… See the full description on the dataset page: https://huggingface.co/datasets/willdaspit/afdb_50_sequence_clustered.tabular100M<n<1B1 likes37 downloads11mo agoHugging Face22cskokgibbs /yeast-tf-sequence-homology-pretokenized-NTtabular1M<n<10M0 likes33 downloads1y agoHugging Face23willdaspit /afdb_50_sequence_clustered_reprstabular10M<n<100M1 likes29 downloads11mo agoHugging Face24ClarusC64 /quantum-gate-sequence-instability-v0.1 quantum-gate-sequence-instability-v0.1 What this dataset does This dataset evaluates whether models can detect instability in quantum gate sequences. Each row represents a simplified quantum circuit execution scenario described through observable device and circuit proxies. The task is to determine whether the gate sequence remains executable inside a stable coherence window or becomes unstable. Core stability idea Quantum gate sequences become unstable when… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-gate-sequence-instability-v0.1.tabulartabular-classificationn<1K0 likes26 downloads5mo agoHugging Face25ClarusC64 /clinical-control-sequence-sepsis-v1 Clinical Control Sequence Sepsis Detection Overview This dataset tests whether a model can detect whether a proposed intervention sequence is the correct stabilizing control sequence for a sepsis-like clinical system. Complex systems are often not stabilized by a single action. They require the correct sequence of interventions delivered in the correct order as the system evolves. The goal of this benchmark is to determine whether the control sequence meaningfully guides… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-control-sequence-sepsis-v1.tabulartext-classificationn<1K0 likes23 downloads6mo agoHugging Face26Prarabdha /Paul_RNA_Sequence_Processed_Datasettabular1K<n<10K1 likes18 downloads3y agoHugging Face27sachithgunasekara /LaMini-instruction-only-SequenceMatcher-Levenstein Dataset Card for "LaMini-LM-filtered-instruction-only-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML" More Information needed tabularn<1K0 likes17 downloads3y agoHugging Face28random-sequence /flock-demo-healthcare-glucose-sectionstabular1K<n<10K0 likes16 downloads8mo agoHugging Face29random-sequence /flock-demo-time-series-prediction-sectionstabular1K<n<10K0 likes16 downloads8mo agoHugging Face30joduor /adaption-ebolavirus-protein-sequences This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-ebolavirus_protein_sequences This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-ebolavirus-protein-sequences.tabular1K<n<10K0 likes15 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.