datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genome-sequence-to-function
AlphaGenome human compact benchmark
This is the materialized human-only 1-Mb benchmark release used by the
genome-sequence-to-function track. It contains 3,200 train, 200 validation,
and 1,000 test examples; all 5,930 human tracks; and eleven output families.
Release ID: release-44d06e9702eba731
Installed size: 92.49 GiB
Input context: 1,048,576 bp
Target span: 196,608 bp
The release keeps the checksummed target objects unchanged and publishes the
current public/label/scoring… See the full description on the dataset page: https://huggingface.co/datasets/zifeng-ai/genome-sequence-to-function.ai4sci-genome-sequence-to-function-data
genome-sequence-to-function: prepared release v4
This is the data for the genome-sequence-to-function task of
T0-RSI/ai4sci-tasks: one model that maps 1,048,576 bp
of human DNA to 5,930 functional-genomics tracks in eleven families, and predicts variant effects.
It is built from AlphaGenome's public training data (gs://alphagenome-datasets/v1/train,
FOLD_0) by the task's environment/data/materialize.sh. Use that script to install it; it
pins this repository's revision and… See the full description on the dataset page: https://huggingface.co/datasets/sunweiwei/ai4sci-genome-sequence-to-function-data.
