datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dna_rendering_processed
DNA-Rendering-Processed Dataset
Project Page | Paper | Code | Model
To enable Diffuman4D model training, we meticulously process the DNA-Rendering dataset by recalibrating camera parameters, optimizing image color correction matrices (CCMs), predicting foreground masks, and estimating human skeletons.
To promote future research in the field of human-centric 3D/4D generation, we have open-sourced our re-annotated labels for the DNA-Rendering dataset in this repo, which includes… See the full description on the dataset page: https://huggingface.co/datasets/krahets/dna_rendering_processed.zoonomia-v1-v3_ccre_non_promoter
bolinas-dna/zoonomia-v1-v3_ccre_non_promoter
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ccre_non_promoter by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (ccre_non_promoter)
ENCODE cCRE V4 non-promoter classes — cre_class != "PLS" (so: dELS, pELS, CA, CA-CTCF, CA-TF, CA-H3K4me3, TF), extended by 500 bp on each side. PLS is excluded because… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_ccre_non_promoter.genomes-v5-genome_set-animals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128
Animals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
242,334,716 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.genomes-v5-genome_set-animals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128
Animals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
68,286,166 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.genomes-v5-genome_set-animals_order204-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128
204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
main). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
101,114,252 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.genomes-v4-genome_set-animals-intervals-v5_256_128genomes-v4-genome_set-animals-intervals-v10_256_128genomes-v4-genome_set-animals-intervals-v11_256_128genomes-v4-genome_set-animals-intervals-v12_256_128zoonomia-v1-v4_ccre_noexon
bolinas-dna/zoonomia-v1-v4_ccre_noexon
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is the v4 ccre_non_promoter arm with every window that overlaps any other functional element (CDS / 3′UTR / ncRNA exon / TSS+5′UTR) removed — i.e. windows whose functional content… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon.zoonomia-v1-v3_cds
bolinas-dna/zoonomia-v1-v3_cds
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled cds by the
snakemake/zoonomia_projection_dataset pipeline
(commit 2ab868a2f1d4).
Region label (cds)
Coding sequence — Ensembl r115 CDS features (get_cds). Highest-priority class: any anchor with overlap on a CDS feature (and union-of-functional fraction ≥ 0.20 across all five labels) is labelled cds, regardless… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v3_cds.genomes-v4-genome_set-animals-intervals-v14_256_128zoonomia-v1-v4_ccre_noexon_enhancer
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.genomes-v4-genome_set-animals-intervals-v13_256_128genomes-v4-genome_set-animals-intervals-v1_256_128DNA_Gen
Citation
Please cite our work using the bibtex below:
BibTeX:
@article{su2025language,
title={Language Models for Controllable DNA Sequence Design},
author={Su, Xingyu and Li, Xiner and Lin, Yuchao and Xie, Ziqian and Zhi, Degui and Ji, Shuiwang},
journal={arXiv preprint arXiv:2507.19523},
year={2025}
}
zoonomia-v1-v1
zoonomia-v1-v1
255 bp human-anchored windows, conservation-filtered (phyloP_447m, proportion_conserved >= 0.20), projected onto 108 family-deduped Zoonomia 447-mammalian assemblies via halLiftover, midpoint-resized to 255 bp, reverse-complement-augmented. Single train split.
Produced by the zoonomia_projection_dataset pipeline in Open-Athena/bolinas-dna — permalinked at the exact code that built this dataset: snakemake/zoonomia_projection_dataset @ 7ff07cd (PR #158).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v1.genomes-v4-genome_set-animals-intervals-v4_512_256dna-format-zoo
DNA Format Zoo
DNA Format Zoo is a collection of public genomic files in varied formats,
encodings, reference builds, and producer-specific dialects. It is intended for
testing parsers, converters, validators, and other bioinformatics tooling against
real files with recorded provenance and expected characteristics.
The initial collection emphasizes HG002 (also known as NA24385), with additional
publicly shared participants included when they provide useful format coverage.… See the full description on the dataset page: https://huggingface.co/datasets/prescience-bio/dna-format-zoo.laya-bio
Laya-Bio: short-sequence candidate-scoring benchmark and reproducibility data
This repository packages the data and saved results used by Laya-Bio: Candidate Scoring and Reliability on Short Biological Sequences (Liang Wang, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology). The main study uses two closed-set tasks, four model conditions and three training seeds, with no additional neural continual pretraining.
Companion… See the full description on the dataset page: https://huggingface.co/datasets/dnagpt/laya-bio.human_ref_dna
Dataset Card for "human_ref_dna"
More Information needed
genomes-v4-genome_set-vertebrates-intervals-v1_255_128-id1_cov1genomes-v4-genome_set-animals-intervals-v6_256_128genomes-v5-genome_set-vertebrates-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128
Vertebrates CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
172,025,122 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-vertebrates-intervals-v5_255_128.riemann-clock-spectra
Riemann Clock spectra
This dataset accompanies the exploratory spectroscopy study in
maris205/riemann_clock, with analysis
base commit 321ae1cd48d69fe4d1091326167a8aaa65f2e597.
It mirrors public, previously reduced, extracted, or coadded quasar spectra and
the study's processed arrays. It contains no newly acquired observations and no
raw CCD frames. The project directory name data/raw/ means downloaded inputs.
This exploratory study reports no detection of Riemann-clock physics… See the full description on the dataset page: https://huggingface.co/datasets/dnagpt/riemann-clock-spectra.genomes-v4-genome_set-vertebrates-intervals-v5_256_128vertebrate-v1-ccre_non_promoter
marin-dna/vertebrate-v1-ccre_non_promoter
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the ccre_non_promoter region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-ccre_non_promoter.vertebrate-v1-issue473-center1-cds
marin-dna/vertebrate-v1-issue473-center1-cds
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
cds cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed #417 protein-coding CDS anchor catalog.
The source projection was produced by the… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-cds.genomes-v4-genome_set-vertebrates-intervals-v1_255_128-id0.3_cov0.3zoonomia-v1-v4_ccre_non_promoter
bolinas-dna/zoonomia-v1-v4_ccre_non_promoter
Per-anchor region-type partition of the cross-mammal training set
bolinas-dna/zoonomia-v1-v1,
restricted to anchors labelled ccre_non_promoter by the v4 region labeler
(snakemake/zoonomia_projection_dataset pipeline,
commit 4729d06d576f).
v4 re-derives the v3 partition with the labeling scheme resolved in
issue #221: base-pair
priority + window majority, a protein-coding-only TSS band, and
ccre_flank=0. See the pipeline README's "v4… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_non_promoter.
