Team Ai
Datasetpublic

sunweiwei/ai4sci-genome-sequence-to-function-data

genome-sequence-to-function: prepared release v4 This is the data for the genome-sequence-to-function task of T0-RSI/ai4sci-tasks: one model that maps 1,048,576 bp of human DNA to 5,930 functional-genomics tracks in eleven families, and predicts variant effects. It is built from AlphaGenome's public training data (gs://alphagenome-datasets/v1/train, FOLD_0) by the task's environment/data/materialize.sh. Use that script to install it; it pins this repository's revision and… See the full description on the dataset page: https://huggingface.co/datasets/sunweiwei/ai4sci-genome-sequence-to-function-data.

sourceHugging Facecc-by-4.0updated 5d agoView on Hugging Face
0likes88downloads
Dataset Card

genome-sequence-to-function: prepared release v4

This is the data for the genome-sequence-to-function task of T0-RSI/ai4sci-tasks: one model that maps 1,048,576 bp of human DNA to 5,930 functional-genomics tracks in eleven families, and predicts variant effects. It is built from AlphaGenome's public training data (gs://alphagenome-datasets/v1/train, FOLD_0) by the task's environment/data/materialize.sh. Use that script to install it; it pins this repository's revision and verifies every archive.

partcontentstored
train targetsall 41,020 eligible intervals of folds 2–7, 8-bit codes of asinh(x) with per-track ranges308 GiB
validation / test targets200 (fold 0) / 1,000 (fold 1) intervals, full precision4 / 21 GiB
supportmanifests, labels, junction scoring queries, track metadata, GRCh38.p13 FASTA, variant evaluations (test queries and labels, labelled development split)4 GiB

Layout.

  • —archives/support.tar holds every file except the targets.
  • —archives/train/*.tar holds one tar per Zarr train shard.
  • —archives/{valid,test}/*.zarr.zip are Zarr ZipStores.
  • —release-manifest.json gives the size and SHA-256 of each of these.

Extracting everything into one directory recreates the release that the task's validator checks.

Test data. The test targets and the test variant labels are included because the upstream data are public. Harbor mounts them only into the task's grader, never into the agent's workspace.

Sources and licenses.

  • —AlphaGenome training data and evaluation tables: Avsec et al., Nature 649 (2026), CC BY 4.0.
  • —TraitGym (Benegas et al. 2025): MIT.
  • —GENCODE v46.
  • —GRCh38.p13 reference genome (GENCODE).

Cite the original works when using this release.