UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model
Rustin's Super Mega Awesome VEDU Model A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive winter-annual grass) across Montana from satellite + environmental data. Science reference: docs/VEDU_48_predictors_detailed.md Data decisions & gotchas: docs/CONTRADICTIONS.md Parity with the Earth Engine build: docs/GEE_PARITY.md Continue-the-build guide: docs/HANDOFF.md Label inventory: docs/DATA_SOURCES.md What it produces 57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.
Rustin's Super Mega Awesome VEDU Model
A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive winter-annual grass) across Montana from satellite + environmental data.
- Science reference: `docs/VEDU_48_predictors_detailed.md`
- Data decisions & gotchas: `docs/CONTRADICTIONS.md`
- Parity with the Earth Engine build: `docs/GEE_PARITY.md`
- Continue-the-build guide: `docs/HANDOFF.md`
- Label inventory: `docs/DATA_SOURCES.md`
What it produces
57 hybrid predictors per plot — 25 phenology (harmonic/Fourier NDVI curve, accumulated heat, and EOS-support features), 25 PRISM/gridMET climate, 4 SSURGO soils, 3 terrain — fed to XGBoost (RF-style) as three model views (phenology 25 · bioclimatic 32 · hybrid 57), evaluated with AUC / PR-AUC / sensitivity / specificity / kappa, and an honest external held-out test on USGS 2025.
A fourth, optional view — `spectral` (34) from Stage 04b — adds seasonal SWIR and red-edge composites. It exists because greenness is the least informative thing in this data: NDVI-only seasonal windows score 0.549 out-of-fold against SWIR's 0.605 and red edge's 0.715. It also transfers best to the external USGS benchmark (0.577, the only view above chance there). See `stages/04b_spectral/`.
Stage 04 can also be re-run as a phenology imagery-source experiment: HLS combined, Landsat-only HLSL30, and Sentinel-only HLSS30. Stages 05-08 accept the same --imagery-source flag, and `stages/11_imagery_comparison/` rolls up the three-way comparison.
Read this before quoting a number
Held-out test AUC is ~0.81–0.82. That number is dominated by comparisons between plots kilometres apart. Score instead the fraction of matched (infested, clean) pairs inside one 5 km block that the model ranks correctly, and skill collapses with distance:
(bioclimatic, out-of-fold, measured-cover labels, mean over 2 fold draws × 3 seeds.)
At paddock scale this model is close to a coin flip; its headline score comes from knowing roughly where a plot is, not what is growing in it. Report both numbers or neither — Stage 07 prints the table by default. Never report bare accuracy either: prevalence is 0.17, so "nothing is infested anywhere" scores 83%.
Stage graph
00_labels ──┬──────────────────────────────────────────────┐
│ (plot_id, lon/lat, obs_year, cover, presence) │
▼ ▼
01_soils 02_climate 03_terrain 04_phenology 04b_spectral 05_compile ──► 06_model
(4) (25) (3) (25) (34) join → 57 XGBoost
│ cGDD ─────────► (GDD-at-transition) optional + spectral 4 views
└────────────────────────────────────────────► + external
USGS test
07_testing · 08_source_transfer · 09_spatial_uncertainty · 10_survey_design
└──────────────► 11_imagery_comparisonEach stage caches by content hash, logs to stages/NN/logs/, and appends to the manifest. 04_phenology consumes 02_climate's daily cGDD series; 05_compile joins the feature stages on plot_id and emits the views + the train/external/CV splits. `04b_spectral` is optional.
Layout
config/ config.yaml · sources.yaml · predictors.yaml (all tunables)
src/vedu/ shared library (cache, logging, manifest, runlog, io, qa, config, paths)
stages/ 00_labels … 10_survey_design, including 09_spatial_uncertainty
(run.py · steps/ · README · inspect.ipynb · outputs · logs · tests)
data/ cache/ (content-addressed) · interim/ · processed/ (manifest.jsonl)
docs/ predictor spec · contradictions · handoff · data sources
run_logs/ runs.jsonl + runs.md (every train/test run, config + metrics)
tests/ library smoke testsQuickstart
# 1) Library sanity + stage status — NO installs needed (src/ is put on sys.path):
PYTHONPATH=src python -m vedu.manifest status
PYTHONPATH=src pytest tests
# 2) Stage 00 (labels) runs on the current env:
python stages/00_labels/run.py # add --force to ignore cache, --dry-run to plan
# Ming is currently enabled in config/sources.yaml; this writes the with-Ming labels.
# 3) Before the remote-data stages, install the rest:
pip install -r requirements.txt # pyarrow, lmfit, netCDF4, scikit-learn, xgboost, joblib
python stages/01_soils/run.py
# ... 02_climate, 03_terrain, 04_phenology, 05_compile, 06_model
# 4) Optional phenology imagery-source comparison:
python stages/04_phenology/run.py --imagery-source landsat --extract-only
python stages/04_phenology/run.py --imagery-source landsat --fit-only
python stages/05_compile/run.py --imagery-source landsat
python stages/06_model/run.py --imagery-source landsat --views phenology,hybrid
python stages/07_testing/run.py --imagery-source landsat --view phenology
python stages/04_phenology/run.py --imagery-source sentinel --extract-only
python stages/04_phenology/run.py --imagery-source sentinel --fit-only
python stages/05_compile/run.py --imagery-source sentinel
python stages/06_model/run.py --imagery-source sentinel --views phenology,hybrid
python stages/07_testing/run.py --imagery-source sentinel --view phenology
python stages/11_imagery_comparison/run.pyDesign decisions (see CONTRADICTIONS.md for the full list)
- Per-observation-year imagery — each plot uses HLS from its own survey year (2023/24/25).
- USGS 2025 is an external held-out test — never trained on; scored separately as the honest benchmark, alongside the pooled random test + spatial CV on the other sources.
- Presence = cover ≥ 20% (heavy-vs-light), trace-VEDU kept as negatives.
- Reproducibility: rerun any stage → cache hits in seconds unless an input, a param, or a step's
CODE_VERSIONchanged.python -m vedu.manifest statusshows what is current vs stale. - Ming data is currently enabled: Stage 00 includes Ming Crow points plus Extra BR point/polygon representatives. Crow polygons stay disabled to avoid double-counting paired Crow observations. To return to the old primary-only dataset, set the three enabled Ming source entries to
enabled: falseinconfig/sources.yaml.
Status
Foundation complete (library + config + docs). Stages build one at a time — see `docs/HANDOFF.md` build-status checklist.
