Team Ai
Datasetpublic

risenfromashes/qraft-experiments

QRAFT experiments Data and results of the experiments on QRAFT, a species-tree method based on quartets, and of its comparison with ASTRAL-X and STELAR-X: simulated data sets, the runs of the three methods, the ablation of QRAFT's pipeline, timings measured with a processor socket to themselves, and runs on biological data. This repository was named risenfromashes/simphy-1M until 2026-10-09; logs written before then still use that name. Folder Contents raw/… See the full description on the dataset page: https://huggingface.co/datasets/risenfromashes/qraft-experiments.

sourceHugging Faceupdated 9h agoView on Hugging Face
0likes2.7kdownloads
Dataset Card

QRAFT experiments

Data and results of the experiments on QRAFT, a species-tree method based on quartets, and of its comparison with ASTRAL-X and STELAR-X: simulated data sets, the runs of the three methods, the ablation of QRAFT's pipeline, timings measured with a processor socket to themselves, and runs on biological data. This repository was named risenfromashes/simphy-1M until 2026-10-09; logs written before then still use that name.

FolderContents
raw/, provenance/SimPhy data sets for the sweep to 1,000,000 taxa and gene trees, 5 replicates, and how they were simulated
vary-ils/data sets at four levels of incomplete lineage sorting (ILS), and QRAFT's runs on them
runs/texas-2026-10-02/the run box of the five-replicate sweep: QRAFT, ASTRAL-X and STELAR-X runs
clean-timing/2026-10-09/timings measured with a socket to themselves
ablation/the ablations of QRAFT's pipeline
angio/, avian-48/, avian-363/, udance/runs on biological data

QRAFT is qraft-38d32f1, the static Release build of commit 38d32f1, unless a measurement build is named. Data set names follow the benchmark harness phylo-bench: t_<taxa>_g_<genes>_sb_<speciation rate>_spmin_<min Ne>_spmax_<max Ne>, with replicates R1–R5.

raw/: the sweep to 1,000,000 taxa and gene trees

Simulated species trees and gene trees for two sweeps used to benchmark species-tree inference at scale:

  • —taxa sweep: 100,000 to 1,000,000 taxa in steps of 100,000, each with 1,000 gene trees;
  • —gene sweep: 100,000 to 1,000,000 gene trees in steps of 100,000, each with 1,000 taxa.

Every size has 5 replicates (R1–R5), each a species tree (s_tree.trees) and its gene trees (all_gt.tre, one Newick tree per line).

Model and command

The data were generated by the benchmark harness phylo-bench (commit 7aa3c87e49bfa6b6032ff81759b21740327385fb), unchanged, with the simphy-x engine in compat mode, which produces trees byte-identical to SimPhy 1.0.2. For every size T taxa × G genes:

bash
./scripts/sim.sh -rs 5 --simphy-data-dir DATA --seed 42 --engine simphy-x-compat \
    --sim-threads all -t T -g G --sb 0.000001 --spmin 100000 --spmax 200000

which runs (see provenance/simphy-commands.txt for each data set's recorded command line):

simphy-x -XM compat -sb f:0.000001 -ld f:0 -lb f:0 -lt f:0 -rs 5 -rl f:G -rg 1 -o OUT \
    -sp u:100000,200000 -su ln:-17.27461,0.6931472 -sg f:1 -sl f:T -st ln:16.2,1 \
    -om 1 -v 2 -od 1 -op 1 -oc 1 -on 1 -cs 42

then concatenates each replicate's gene trees into all_gt.tre (dropping SimPhy's _0_0 label suffixes) with the harness's concat_gene_trees.py and reorganize_trees.py. Speciation rate 1e-6, no duplication, loss or transfer, effective population size uniform in [100,000, 200,000], species-tree height lognormal(16.2, 1), seed 42.

The 100k, 200k and 300k data sets of both sweeps are identical to those of the ASTRAL-X benchmark (imAniksahA/blab, ph/d/simulated/astralx-datasets/raw/): every tree file has the same CRC-32 and size. Only SimPhy's .command, .params and .db files differ, since they record the binary's and the output's paths.

Layout

raw/t_T_g_G_sb_0.000001_spmin_100000_spmax_200000.zip           (T × G ≤ 0.6e9)
raw/t_T_g_G_sb_0.000001_spmin_100000_spmax_200000.R1-R3.zip     (T × G ≥ 0.7e9: two parts,
raw/t_T_g_G_sb_0.000001_spmin_100000_spmax_200000.R4-R5.zip      under the 50 GB file limit)
provenance/

Each zip holds the data set's directory with SimPhy's .command, .params and .db files and its replicates, R<n>/all_gt.tre and R<n>/s_tree.trees. Unzip both parts of a split data set into the same directory to get R1–R5:

bash
unzip t_1000000_g_1000_sb_0.000001_spmin_100000_spmax_200000.R1-R3.zip
unzip -o t_1000000_g_1000_sb_0.000001_spmin_100000_spmax_200000.R4-R5.zip

Check the trees against provenance/sha256-trees.txt (sha256sum -c after cd into the extraction directory).

Provenance

  • —sim-commands.log: every command run (sim.sh and zip), verbatim, with start and end time, working directory and exit status (JSON lines).
  • —simphy-commands.txt: the SimPhy command line each data set recorded.
  • —sha256-trees.txt, sha256-zips.txt: sha256 of every tree file and every zip (for a zip packed more than once, the last line is the uploaded one).
  • —environment-helper*.txt: the machines, the engine version and the sha256 of the harness scripts (identical on every machine).
  • —NOTES.txt: what happened during the run, in order.
  • —sim1m.py (driver), zip-parallel.py (multi-core zip -r), reorg_guard.py (memory guard), cleanup_restart.sh, pack_manual.py: the scripts that ran around the harness.
  • —logs/helper<n>-*.tgz: each machine's harness simulation logs and driver logs.

The data sets were simulated on rented vast.ai machines, with the same harness scripts (same sha256) on each:

MachineData sets
helper 1: Ryzen 9 9950X, Minnesotataxa 100k, 200k, 300k, 500k–900k; genes 100k, 800k
helper 2: Ryzen 9 9950X, Japannone (its disk stalled; abandoned)
helper 3: EPYC 7B13, Nebraskataxa 1M; genes 900k, 1M
helper 4: Ryzen 9 9950X, Utahtaxa 400k; genes 200k–700k

Helper 1 rebooted twice, and helper 3's filesystem stalled while deleting SimPhy's per-locus files, after its three data sets were complete. Data sets in progress at a failure were simulated again from the start; completed ones were kept. NOTES.txt records every step. Zips were made with Info-ZIP 3.0 or, later, with zip-parallel.py: same layout and deflate format, slightly different bytes; the tree files are identical either way.

vary-ils/: four levels of ILS

The varying-ILS benchmark of phylo-bench: 40 SimPhy data sets with five replicates each. They are 1,000 to 40,000 taxa with 1,000 gene trees, and 1,000 taxa with 5,000 to 50,000 gene trees, each at four ILS levels set by the largest effective population size: spmax 150,000, 200,000, 250,000 and 300,000 for ILS-L1 to ILS-L4 (spmin 100,000, speciation rate 1e-6). ILS-L2 has the sweep's settings.

  • —raw/<data set>.zip: the data sets, laid out as in the top-level raw/.
  • —qraft_outputs/<data set>/: QRAFT with 64 threads on every replicate.
  • —stat-vary-ils-qraft.csv and stat-vary-ils-qraft-summary.csv: one row per run, and the mean and standard deviation per data set.
  • —run-vary-ils.log, ils.log, collect.log: the harness's logs.

116 of these 200 runs shared their socket with ablation jobs, so their times are too long: 30,000 taxa at ILS-L1 R5 and at ILS-L2 to L4, all of 40,000 taxa, and 1,000 taxa with 5,000 to 50,000 genes. clean-timing/2026-10-09/vils-clean.csv holds their clean times. The trees, and so the RF rates, do not depend on the load.

runs/texas-2026-10-02/: the five-replicate sweep

The vast.ai machine the sweep ran on from 2026-10-02 to 2026-10-09: two 32-core AMD EPYC processors, 409 GiB of memory and an RTX 4090. Every run had one socket to itself, with 64 threads. QRAFT has up to five replicates per size, and three for each of 700,000 to 1,000,000 taxa. ASTRAL-X and STELAR-X ran as references, mostly one replicate per size; results/runs.csv lists every run and its status. Sizes not in raw/, the finer steps up to 300,000, come from the ASTRAL-X benchmark's data (imAniksahA/blab).

  • —results/runs.csv and results/jobs.jsonl: every run, with its status, wall time, peak memory and RF rate. results/runbox.log is the driver's log, results/logs/ holds the harness logs, results/environment.txt describes the machine and results/thermal.log its clocks and temperatures.
  • —results/r275k/: the third 275,000-taxon replicate, R2, run after the driver had finished.
  • —results/highils/, results/bio/, results/udance/: high-ILS replicates rerun without a time limit, and the biological and uDance runs of this machine.
  • —results/CONTAMINATED.txt: one ASTRAL-X run whose time is not usable.
  • —outputs/: the harness's outputs (qraft_outputs/, astralx_outputs/, stelarx_outputs/), with trees, statistics and command records.
  • —health/: the GPU and QRAFT checks run when the machine was rented.

clean-timing/2026-10-09/: clean timings

Runs timed with a socket to themselves and nothing else on it, by clean_runner.sh (one runner per socket) from the queue that build_queue.py wrote. queue/done/ holds each item's script and log.

  • —vils-clean.csv (the harness's columns, no header): the 116 varying-ILS runs above, rerun alone.
  • —timing.csv (group,dataset,rep,variant,wall_s,maxrss_kb,exit):
  • —the nine pipeline variants on ASTRAL-II, 1,000-taxon, ASTRAL-MP 10k, ASTRAL-III, 1KP, B10K exons and intergenic, the angiosperms and avian-48. The variants are qraft, no_nni, mre, mre_nni, mre_nni_div, balanced, no_test, regraft_norm and nni_to_end. They ran with the measurement build qraft-nni-test3, whose default output equals the release's.
  • —true against estimated gene trees at 500, 1,000 and 10,000 taxa;
  • —the biological data sets, three runs each.
  • —discordance.csv (no header; data set, replicate, taxa, genes, genes sampled, GT-ST %, gene-tree pairs sampled, GT-GT %): the discordance of gene trees with the species tree (GT-ST) and between pairs of gene trees (GT-GT), computed by discord.py.
  • —results/: the trees and the time and memory record of every run.

Five varying-ILS replicates took 736 to 7,516 s where their siblings took about 100 to 180 s: 30,000 taxa at L1 R5, L3 R4 and L4 R3, and 40,000 taxa at L1 R5 and L2 R5. Each has far more discordant gene trees than its siblings.

ablation/: the ablations of QRAFT's pipeline

  • —2026-10-08/: the full ablation, 1,706 jobs with none failed. ablation-all.csv has one row per tree, 45,452 rows with RF and the tqDist quartet score. trees/<family>.tar.gz holds each family's trees. The families are the 37-, 48-, 100-, 200-, 500- and 1,000-taxon sets, ASTRAL-II, ASTRAL-III 2000, ASTRAL-MP 10k, 1KP and the varying-ILS data sets.
  • —Bases: the default pipeline, MRE only, balanced divide, no bipartition test, normalized regraft, 4 proposals with 2 anchor pairs, and MRE, then NNI, then divide.
  • —Each base runs without NNI and with three NNI modes: the release's (at most 20 moves, support gate), batched with no move cap, and batched with no cap but with the gate.
  • —The varying-ILS sets have only the default and MRE bases.
  • —The times are not clean; clean-timing/ has clean ones.
  • —2026-09-30-texas/ and 2026-10-01-local/: the first ablation, on an earlier machine, and the quartet scores computed locally afterwards. Their READMEs describe them.

Biological data

QRAFT with 64 threads on four data sets. Each inputs.tsv names the input files, which are not copied here, and each biological-dataset.command says where they come from.

  • —angio/: the angiosperms (1KP), 9,524 taxa, 353 gene trees.
  • —avian-48/: 48 birds, 14,446 gene trees.
  • —avian-363/: B10K intergenic, 363 birds, 63,430 gene trees.
  • —It also holds an ASTRAL-X run and two QRAFT trees with uncapped NNI.
  • —qraft-vs-astralx.txt compares their time, memory, RF and quartet scores.
  • —udance/qraft_outputs/batch-nni/: QRAFT end to end on the pruned uDance gene trees (380 genes, 199,330 taxa), with batched NNI and no move cap, run with the measurement build. README.txt gives the settings and nni.trace every accepted move.
risenfromashes/qraft-experiments · Team Ai