Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01just-dna-seq /annotators Genomic Variant Annotators Curated genomic variant annotation modules from the DNA-seq project. Overview This dataset contains pre-computed annotation data for genetic variants, organized by module: Module Description Files longevitymap Longevity-associated variants annotations.parquet, studies.parquet, weights.parquet Schema annotations.parquet Variant-level facts linking rsIDs to genes and phenotypes. rsid: dbSNP… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/annotators.image1K<n<10K0 likes4.3k downloads13d agoHugging Face02krahets /dna_rendering_processedgated DNA-Rendering-Processed Dataset Project Page | Paper | Code | Model To enable Diffuman4D model training, we meticulously process the DNA-Rendering dataset by recalibrating camera parameters, optimizing image color correction matrices (CCMs), predicting foreground masks, and estimating human skeletons. To promote future research in the field of human-centric 3D/4D generation, we have open-sourced our re-annotated labels for the DNA-Rendering dataset in this repo, which includes… See the full description on the dataset page: https://huggingface.co/datasets/krahets/dna_rendering_processed.imageimage-to-3d1K<n<10K9 likes1.3k downloads11mo agoHugging Face03marin-dna /zoonomia-v1-v3_ccre_non_promoter bolinas-dna/zoonomia-v1-v3_ccre_non_promoter Per-anchor region-type partition of the cross-mammal training set bolinas-dna/zoonomia-v1-v1, restricted to anchors labelled ccre_non_promoter by the snakemake/zoonomia_projection_dataset pipeline (commit 2ab868a2f1d4). Region label (ccre_non_promoter) ENCODE cCRE V4 non-promoter classes — cre_class != "PLS" (so: dELS, pELS, CA, CA-CTCF, CA-TF, CA-H3K4me3, TF), extended by 500 bp on each side. PLS is excluded because… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ccre_non_promoter.tabular10M<n<100M0 likes1.3k downloads1mo agoHugging Face04marin-dna /genomes-v5-genome_set-animals-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128 Animals CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 242,334,716 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.text100M<n<1B0 likes1.3k downloads1mo agoHugging Face05marin-dna /genomes-v5-genome_set-animals-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128 Animals promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 68,286,166 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.text10M<n<100M0 likes1.3k downloads1mo agoHugging Face06marin-dna /genomes-v5-genome_set-animals_order204-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128 204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit main). Each repo in the family is one (genome_set, region-recipe) combination. Size 101,114,252 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.text100M<n<1B0 likes1.2k downloads3mo agoHugging Face07marin-dna /genomes-v4-genome_set-animals-intervals-v5_256_128text100M<n<1B0 likes1.2k downloads9mo agoHugging Face08marin-dna /genomes-v4-genome_set-animals-intervals-v10_256_128text100M<n<1B0 likes1.1k downloads8mo agoHugging Face09marin-dna /genomes-v4-genome_set-animals-intervals-v11_256_128text100M<n<1B0 likes1.1k downloads8mo agoHugging Face10marin-dna /genomes-v4-genome_set-animals-intervals-v12_256_128text100M<n<1B0 likes1.1k downloads8mo agoHugging Face11marin-dna /genomes-v4-genome_set-animals-intervals-v7_256_1280 likes1.1k downloads9mo agoHugging Face12marin-dna /zoonomia-v1-v4_ccre_noexon bolinas-dna/zoonomia-v1-v4_ccre_noexon A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is the v4 ccre_non_promoter arm with every window that overlaps any other functional element (CDS / 3′UTR / ncRNA exon / TSS+5′UTR) removed — i.e. windows whose functional content… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon.tabular10M<n<100M0 likes1k downloads3mo agoHugging Face13marin-dna /zoonomia-v1-v3_cds bolinas-dna/zoonomia-v1-v3_cds Per-anchor region-type partition of the cross-mammal training set bolinas-dna/zoonomia-v1-v1, restricted to anchors labelled cds by the snakemake/zoonomia_projection_dataset pipeline (commit 2ab868a2f1d4). Region label (cds) Coding sequence — Ensembl r115 CDS features (get_cds). Highest-priority class: any anchor with overlap on a CDS feature (and union-of-functional fraction ≥ 0.20 across all five labels) is labelled cds, regardless… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_cds.tabular10M<n<100M0 likes1k downloads5mo agoHugging Face14marin-dna /genomes-v4-genome_set-animals-intervals-v14_256_128text10M<n<100M0 likes1k downloads8mo agoHugging Face15marin-dna /zoonomia-v1-v4_ccre_noexon_enhancer bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.tabular10M<n<100M0 likes1k downloads3mo agoHugging Face16marin-dna /genomes-v4-genome_set-animals-intervals-v13_256_128text10M<n<100M0 likes1k downloads8mo agoHugging Face17marin-dna /genomes-v4-genome_set-animals-intervals-v1_256_128text10M<n<100M0 likes998 downloads9mo agoHugging Face18xingyusu /DNA_Gen Citation Please cite our work using the bibtex below: BibTeX: @article{su2025language, title={Language Models for Controllable DNA Sequence Design}, author={Su, Xingyu and Li, Xiner and Lin, Yuchao and Xie, Ziqian and Zhi, Degui and Ji, Shuiwang}, journal={arXiv preprint arXiv:2507.19523}, year={2025} } document10K<n<100K3 likes976 downloads1y agoHugging Face19marin-dna /zoonomia-v1-v1 zoonomia-v1-v1 255 bp human-anchored windows, conservation-filtered (phyloP_447m, proportion_conserved >= 0.20), projected onto 108 family-deduped Zoonomia 447-mammalian assemblies via halLiftover, midpoint-resized to 255 bp, reverse-complement-augmented. Single train split. Produced by the zoonomia_projection_dataset pipeline in Open-Athena/bolinas-dna — permalinked at the exact code that built this dataset: snakemake/zoonomia_projection_dataset @ 7ff07cd (PR #158).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v1.tabular100M<n<1B0 likes953 downloads5mo agoHugging Face20marin-dna /genomes-v4-genome_set-animals-intervals-v4_512_256text10M<n<100M0 likes887 downloads9mo agoHugging Face21prescience-bio /dna-format-zoo DNA Format Zoo DNA Format Zoo is a collection of public genomic files in varied formats, encodings, reference builds, and producer-specific dialects. It is intended for testing parsers, converters, validators, and other bioinformatics tooling against real files with recorded provenance and expected characteristics. The initial collection emphasizes HG002 (also known as NA24385), with additional publicly shared participants included when they provide useful format coverage.… See the full description on the dataset page: https://huggingface.co/datasets/prescience-bio/dna-format-zoo.text10K<n<100K0 likes826 downloads4d agoHugging Face22dnagpt /laya-bio Laya-Bio: short-sequence candidate-scoring benchmark and reproducibility data This repository packages the data and saved results used by Laya-Bio: Candidate Scoring and Reliability on Short Biological Sequences (Liang Wang, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology). The main study uses two closed-set tasks, four model conditions and three training seeds, with no additional neural continual pretraining. Companion… See the full description on the dataset page: https://huggingface.co/datasets/dnagpt/laya-bio.texttext-classification100K<n<1M1 likes801 downloads17d agoHugging Face23nuvocare /human_ref_dna Dataset Card for "human_ref_dna" More Information needed tabular1M<n<10M1 likes732 downloads3y agoHugging Face24marin-dna /genomes-v4-genome_set-vertebrates-intervals-v1_255_128-id1_cov1text10M<n<100M0 likes698 downloads7mo agoHugging Face25marin-dna /genomes-v4-genome_set-animals-intervals-v6_256_128text10M<n<100M0 likes691 downloads9mo agoHugging Face26marin-dna /genomes-v5-genome_set-vertebrates-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128 Vertebrates CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 172,025,122 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128.text100M<n<1B0 likes687 downloads4mo agoHugging Face27dnagpt /riemann-clock-spectra Riemann Clock spectra This dataset accompanies the exploratory spectroscopy study in maris205/riemann_clock, with analysis base commit 321ae1cd48d69fe4d1091326167a8aaa65f2e597. It mirrors public, previously reduced, extracted, or coadded quasar spectra and the study's processed arrays. It contains no newly acquired observations and no raw CCD frames. The project directory name data/raw/ means downloaded inputs. This exploratory study reports no detection of Riemann-clock physics… See the full description on the dataset page: https://huggingface.co/datasets/dnagpt/riemann-clock-spectra.imagen<1K0 likes665 downloads19d agoHugging Face28marin-dna /genomes-v4-genome_set-vertebrates-intervals-v5_256_128text100M<n<1B0 likes657 downloads8mo agoHugging Face29marin-dna /vertebrate-v1-ccre_non_promoter marin-dna/vertebrate-v1-ccre_non_promoter Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the ccre_non_promoter region cohort with all species scope and preserves source FASTA/2bit letter case. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence case is independent of that filter: lowercase bases preserve source repeat masking, uppercase bases… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-ccre_non_promoter.tabular10M<n<100M0 likes650 downloads2mo agoHugging Face30marin-dna /vertebrate-v1-issue473-center1-cds marin-dna/vertebrate-v1-issue473-center1-cds Review status: draft generated for issue #473 review before upload. Human-anchored 255 bp vertebrate sequences for the cds cohort under the center_1 policy. The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed #417 protein-coding CDS anchor catalog. The source projection was produced by the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-cds.tabular10M<n<100M0 likes641 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.