Taykhoom/functional-multiclass-gamba
GAMBA Functional Region Multiclass This representation benchmark asks whether frozen sequence embeddings separate genomic functional categories. Each row is one annotated region; label == category. Loading from datasets import load_dataset full_bidi = load_dataset( "Taykhoom/functional-multiclass-gamba", "full-bidi", split="all", ) paper_test = full_bidi.filter( lambda row: row["split"] == "test" and row["category"] != "noncoding_regions" )… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/functional-multiclass-gamba.
GAMBA Functional Region Multiclass
This representation benchmark asks whether frozen sequence embeddings separate genomic functional categories. Each row is one annotated region; label == category.
Loading
from datasets import load_dataset
full_bidi = load_dataset(
"Taykhoom/functional-multiclass-gamba",
"full-bidi",
split="all",
)
paper_test = full_bidi.filter(
lambda row: row["split"] == "test"
and row["category"] != "noncoding_regions"
)Choosing a configuration
The two dimensions are independent:
The full name refers to the retained ROI visible in the 2,048 bp context; very long features may be clipped according to GAMBA's context rules. It does not mean an unlimited genomic interval.
The 100 bp genomic target is selected once from the source interval before either context view is constructed. Causal and bidirectional rows therefore have identical chrom, start, and end targets and retain stable pair_id values for unchanged source examples.
Dataset size
The corrected paper view excludes category == "noncoding_regions".
Labels
Full configs contain eleven labels:
The 100 bp configs contain 65,655 rows, including 8,942 noncoding targets. Promoter entries do not satisfy the 100 bp criterion and are absent; vista_enhancer has 690 rows, and other categories are filtered by source interval length. Every 100 bp sequence context is 2,048 bp. For full configs, all 92,482 bidirectional contexts are 2,048 bp; causal has 82,502 contexts at 2,048 bp and 9,980 truncated long-feature contexts at 1,000 bp.
Evaluation protocol
For the corrected paper evaluation:
- filter
split == "test"; - exclude
noncoding_regions; - pool embeddings over the stored span;
- perform cosine leave-one-out 1-nearest-neighbor classification;
- report balanced accuracy.
No supervised model-fitting split is implied by Hugging Face split all.
Columns
RepeatMasker correction
RepeatMasker genomic strand and repeat class are stored separately. This release corrects an earlier local conversion that placed the repeat class in the strand column, then regenerates repeat-derived contexts, pools, controls, and phyloP features. Labels are unchanged.
Only VISTA rows with hg38 assembly and normalized positive expression are eligible: 1,367 raw records collapse to 1,240 unique element-coordinate records; 1,522 non-positive hg38 and 1,750 non-hg38 rows are excluded. Inputs containing ambiguous bases are removed globally. Exact inputs with conflicting category labels are excluded as a group; same-label duplicates share context_group_id. No exact input crosses train/test. ATG test inputs are reserved globally; the exact duplicate chr7 functional training row is therefore excluded from full multiclass configs.
Processing and citation
Processing and verification:
Consens, M. E. et al. Predicting evolutionary rate as a pretraining task improves genome language model representations. bioRxiv (2026). https://doi.org/10.64898/2026.02.02.703275
License
The processing code derived from GAMBA is MIT licensed under the processing repository's LICENSE. This generated dataset is marked other: incorporated reference sequence, annotations, and phyloP-derived values retain their upstream terms, so no blanket MIT license is asserted for the Parquets.
