Team Ai
Datasetpublic

Taykhoom/functional-multiclass-gamba

GAMBA Functional Region Multiclass This representation benchmark asks whether frozen sequence embeddings separate genomic functional categories. Each row is one annotated region; label == category. Loading from datasets import load_dataset full_bidi = load_dataset( "Taykhoom/functional-multiclass-gamba", "full-bidi", split="all", ) paper_test = full_bidi.filter( lambda row: row["split"] == "test" and row["category"] != "noncoding_regions" )… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/functional-multiclass-gamba.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes50downloads
Dataset Card

GAMBA Functional Region Multiclass

This representation benchmark asks whether frozen sequence embeddings separate genomic functional categories. Each row is one annotated region; label == category.

Loading

python
from datasets import load_dataset

full_bidi = load_dataset(
    "Taykhoom/functional-multiclass-gamba",
    "full-bidi",
    split="all",
)

paper_test = full_bidi.filter(
    lambda row: row["split"] == "test"
    and row["category"] != "noncoding_regions"
)

Choosing a configuration

The two dimensions are independent:

DimensionOptionsMeaning
Contextcausal, bidiEnd-anchor ROI for autoregressive models or center it for bidirectional models.
Poolingfull, 100bpPool over the retained ROI or a deterministic 100 bp subspan.

The full name refers to the retained ROI visible in the 2,048 bp context; very long features may be clipped according to GAMBA's context rules. It does not mean an unlimited genomic interval.

The 100 bp genomic target is selected once from the source interval before either context view is constructed. Causal and bidirectional rows therefore have identical chrom, start, and end targets and retain stable pair_id values for unchanged source examples.

Dataset size

Config familyComplete rowsHeld-out testCorrected paper rowsCorrected paper test
full-*92,48218,40482,98016,506
100bp-*65,65513,09556,71311,324

The corrected paper view excludes category == "noncoding_regions".

Labels

Full configs contain eleven labels:

CategoryRows
repeats9,987
UCNE4,169
vista_enhancer691
promoters9,961
UTR59,809
UTR39,717
coding_regions9,904
exons9,926
introns9,026
upstream_TSS9,790
noncoding_regions9,502

The 100 bp configs contain 65,655 rows, including 8,942 noncoding targets. Promoter entries do not satisfy the 100 bp criterion and are absent; vista_enhancer has 690 rows, and other categories are filtered by source interval length. Every 100 bp sequence context is 2,048 bp. For full configs, all 92,482 bidirectional contexts are 2,048 bp; causal has 82,502 contexts at 2,048 bp and 9,980 truncated long-feature contexts at 1,000 bp.

Evaluation protocol

For the corrected paper evaluation:

  1. 1.filter split == "test";
  2. 2.exclude noncoding_regions;
  3. 3.pool embeddings over the stored span;
  4. 4.perform cosine leave-one-out 1-nearest-neighbor classification;
  5. 5.report balanced accuracy.

No supervised model-fitting split is implied by Hugging Face split all.

Columns

Column groupDescription
split, sequence, label, pair_idChromosome partition, exact input, functional class, and stable feature ID.
context_group_idExact model-input leakage group within this physical config.
category, repeat_class, scope, context_policySource class, nullable repeat class, full/100bp scope, and context geometry.
chrom, start, end, source_strandZero-based, half-open hg38 coordinates and biological strand; 100 bp configs store the selected target interval.
sequence_orientationExplicit +/- orientation used for sequence geometry and reverse complementation.
context_start, context_endForward-genome context coordinates.
roi_start, roi_endRetained full-feature or selected 100 bp target offsets in oriented sequence.
pool_start_in_window, pool_end_in_windowEvaluation span, identical to the stored ROI.
nameSource annotation identifier.
phylop_*, phylop_context_*Pooling-span and symmetric-context phyloP summaries.

RepeatMasker correction

RepeatMasker genomic strand and repeat class are stored separately. This release corrects an earlier local conversion that placed the repeat class in the strand column, then regenerates repeat-derived contexts, pools, controls, and phyloP features. Labels are unchanged.

Only VISTA rows with hg38 assembly and normalized positive expression are eligible: 1,367 raw records collapse to 1,240 unique element-coordinate records; 1,522 non-positive hg38 and 1,750 non-hg38 rows are excluded. Inputs containing ambiguous bases are removed globally. Exact inputs with conflicting category labels are excluded as a group; same-label duplicates share context_group_id. No exact input crosses train/test. ATG test inputs are reserved globally; the exact duplicate chr7 functional training row is therefore excluded from full multiclass configs.

Processing and citation

Processing and verification:

Consens, M. E. et al. Predicting evolutionary rate as a pretraining task improves genome language model representations. bioRxiv (2026). https://doi.org/10.64898/2026.02.02.703275

License

The processing code derived from GAMBA is MIT licensed under the processing repository's LICENSE. This generated dataset is marked other: incorporated reference sequence, annotations, and phyloP-derived values retain their upstream terms, so no blanket MIT license is asserted for the Parquets.