datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hgt-bootstrap-v1-synthetic
HGT Bootstrap V1 Synthetic Pairs
8,112 synthetic mutation pairs generated via ICI-DC (Interleaved Codon Insertion — Double Consensus) using the SAD coeff1.5 checkpoint as Model A and HyenaDNA-tiny-1k as Model B.
Generation Details
Model A: Nhoodie/omni-dna-sad-mutation (SAD coeff1.5 checkpoint)
Model B: LongSafari/hyenadna-tiny-1k-seqlen-hf (Legacy DC)
Source sequences: 1,014 unique sequences from 8 taxonomic domains
Gap intervals: [3, 4, 5, 6] (codon distance between… See the full description on the dataset page: https://huggingface.co/datasets/Nhoodie/hgt-bootstrap-v1-synthetic.ppl-synthesis-sft-bootstrap
SynthStats PPL Synthesis SFT Bootstrap
This dataset contains natural-language modelling prompts paired with
probabilistic programs, written in the probabilistic programming
languages PyMC (Python) and LazyPPL (Haskell), for supervised
fine-tuning (SFT).
Each row has these fields:
prompt: natural-language modelling task.
reasoning_trace: modelling rationale for the program.
completion: one fenced program block.
complexity: coarse task complexity label.
metadata: runtime… See the full description on the dataset page: https://huggingface.co/datasets/SynthStats/ppl-synthesis-sft-bootstrap.hgt-bootstrap-v2-synthetic
HGT Bootstrap V2 Synthetic Pairs
259,896 synthetic mutation pairs (129,948 train + 129,948 eval) generated via ICI-DC using the bootstrap S2 checkpoint.
Generation Details
Model A: Nhoodie/omni-dna-sad-mutation-bootstrap (S2 bootstrap checkpoint)
Model B: LongSafari/hyenadna-tiny-1k-seqlen-hf
Source sequences: 4,998 unique sequences from 8 taxonomic domains
Gap intervals: [3, 4, 5, 6] (codon distance between gaps)
Seeds: Train=[42, 137, 7, 23, 31, 89, 53], Eval=[101, 157… See the full description on the dataset page: https://huggingface.co/datasets/Nhoodie/hgt-bootstrap-v2-synthetic.Bootstrap5
