SparseWake/sparsewake
SparseWake SparseWake is a synthetic benchmark for sparse temporal hydrodynamic sensing. ICLR 2027 release The expanded release adds controlled multi-source mixtures and common-prior nearest-source tasks, with complete core data banks, reference checkpoints, a small review supplement, and reproduction code with a frozen wake-library input. Download release iclr2027-v1.0rc2 The version page lists the three archives, exact sizes, checksums, extraction instructions… See the full description on the dataset page: https://huggingface.co/datasets/SparseWake/sparsewake.
0275
1# SparseWake Anonymous Release Checklist2 3## Required Before Upload4 5- [ ] Upload only the release-tree contents; do not include the parent `.git` directory.6- [ ] Confirm the three full HDF5 files are present in `data/processed/`.7- [ ] Run `python scripts/verify_dataset.py --data data/sample/sparsewake_sample.h5 --checksums data/checksums.sha256` after final packaging contents are fixed.8- [ ] Confirm `MANIFEST.json` and `data/checksums.sha256` match the final upload contents.9- [ ] Inspect the uploaded file list for hidden files, notebooks, local logs, cache folders, and OS metadata.10- [ ] Do not include development directories outside the artifact tree.11- [ ] Confirm Croissant metadata remains anonymous and does not contain creator email, institution, publisher, non-anonymous URL, or local path.12- [ ] Confirm the license is listed as `cc-by-4.0` in the dataset card.13 14## Reviewer Smoke-Test Commands15 16```bash17python -m venv .venv18.venv/Scripts/pip install -r requirements.txt19python scripts/verify_dataset.py --data data/sample/sparsewake_sample.h5 --checksums data/checksums.sha25620python scripts/train_temporal_mlp.py --config configs/main_v04.yaml --data data/sample/sparsewake_sample.h5 --quick21python scripts/reproduce_tables.py --results data/results --out tables22python scripts/make_main_figures.py --results data/results --out figures23python scripts/make_supp_figures.py --results data/results --out figures24```25 26Use `.venv/bin/pip` and `.venv/bin/python` on Unix-like systems.27 28## Final Manual Review29 30- [ ] README states benchmark purpose, dataset structure, HDF5 fields, targets, sensors, splits, metrics, reproduction commands, compute expectations, limitations, license, and citation.31- [ ] Dataset card states synthetic/simulated origin, intended and out-of-scope use, limitations, no human or private data, and why upstream simulator/source code is not redistributed.32- [ ] Referenced figures exist at `figures/main/Fig1.pdf` through `Fig3.pdf` and `figures/supp/Fig4.pdf` through `Fig9.pdf`.33- [ ] Stored summary CSVs in `data/results/` match manuscript tables and figures within rounding tolerance.34- [ ] `AUDIT_REPORT.md` lists unresolved issues, changed files, tested commands, and revision notes.35 