Team Ai
Datasetpublic

UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model

Rustin's Super Mega Awesome VEDU Model A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive winter-annual grass) across Montana from satellite + environmental data. Science reference: docs/VEDU_48_predictors_detailed.md Data decisions & gotchas: docs/CONTRADICTIONS.md Parity with the Earth Engine build: docs/GEE_PARITY.md Continue-the-build guide: docs/HANDOFF.md Label inventory: docs/DATA_SOURCES.md What it produces 57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.

sourceHugging Faceupdated 22d agoView on Hugging Face
0likes532downloads
Dataset Card

Rustin's Super Mega Awesome VEDU Model

A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive winter-annual grass) across Montana from satellite + environmental data.

  • —Science reference: `docs/VEDU_48_predictors_detailed.md`
  • —Data decisions & gotchas: `docs/CONTRADICTIONS.md`
  • —Parity with the Earth Engine build: `docs/GEE_PARITY.md`
  • —Continue-the-build guide: `docs/HANDOFF.md`
  • —Label inventory: `docs/DATA_SOURCES.md`

What it produces

57 hybrid predictors per plot — 25 phenology (harmonic/Fourier NDVI curve, accumulated heat, and EOS-support features), 25 PRISM/gridMET climate, 4 SSURGO soils, 3 terrain — fed to XGBoost (RF-style) as three model views (phenology 25 · bioclimatic 32 · hybrid 57), evaluated with AUC / PR-AUC / sensitivity / specificity / kappa, and an honest external held-out test on USGS 2025.

A fourth, optional view — `spectral` (34) from Stage 04b — adds seasonal SWIR and red-edge composites. It exists because greenness is the least informative thing in this data: NDVI-only seasonal windows score 0.549 out-of-fold against SWIR's 0.605 and red edge's 0.715. It also transfers best to the external USGS benchmark (0.577, the only view above chance there). See `stages/04b_spectral/`.

Stage 04 can also be re-run as a phenology imagery-source experiment: HLS combined, Landsat-only HLSL30, and Sentinel-only HLSS30. Stages 05-08 accept the same --imagery-source flag, and `stages/11_imagery_comparison/` rolls up the three-way comparison.

Read this before quoting a number

Held-out test AUC is ~0.81–0.82. That number is dominated by comparisons between plots kilometres apart. Score instead the fraction of matched (infested, clean) pairs inside one 5 km block that the model ranks correctly, and skill collapses with distance:

separation< 250 m250–500 m0.5–1 km1–2 km> 2 km
concordance0.5760.5870.6800.7770.867

(bioclimatic, out-of-fold, measured-cover labels, mean over 2 fold draws × 3 seeds.)

At paddock scale this model is close to a coin flip; its headline score comes from knowing roughly where a plot is, not what is growing in it. Report both numbers or neither — Stage 07 prints the table by default. Never report bare accuracy either: prevalence is 0.17, so "nothing is infested anywhere" scores 83%.

Stage graph

00_labels ──┬──────────────────────────────────────────────┐
            │  (plot_id, lon/lat, obs_year, cover, presence) │
            ▼                                                ▼
   01_soils  02_climate  03_terrain  04_phenology  04b_spectral   05_compile ──► 06_model
     (4)        (25)         (3)          (25)          (34)        join → 57       XGBoost
                   │ cGDD ─────────► (GDD-at-transition)  optional   + spectral    4 views
                   └────────────────────────────────────────────►                  + external
                                                                                    USGS test
                 07_testing · 08_source_transfer · 09_spatial_uncertainty · 10_survey_design
                                           └──────────────► 11_imagery_comparison

Each stage caches by content hash, logs to stages/NN/logs/, and appends to the manifest. 04_phenology consumes 02_climate's daily cGDD series; 05_compile joins the feature stages on plot_id and emits the views + the train/external/CV splits. `04b_spectral` is optional.

Layout

config/     config.yaml · sources.yaml · predictors.yaml   (all tunables)
src/vedu/   shared library (cache, logging, manifest, runlog, io, qa, config, paths)
stages/     00_labels … 10_survey_design, including 09_spatial_uncertainty
            (run.py · steps/ · README · inspect.ipynb · outputs · logs · tests)
data/       cache/ (content-addressed) · interim/ · processed/ (manifest.jsonl)
docs/       predictor spec · contradictions · handoff · data sources
run_logs/   runs.jsonl + runs.md  (every train/test run, config + metrics)
tests/      library smoke tests

Quickstart

bash
# 1) Library sanity + stage status — NO installs needed (src/ is put on sys.path):
PYTHONPATH=src python -m vedu.manifest status
PYTHONPATH=src pytest tests

# 2) Stage 00 (labels) runs on the current env:
python stages/00_labels/run.py            # add --force to ignore cache, --dry-run to plan

# Ming is currently enabled in config/sources.yaml; this writes the with-Ming labels.

# 3) Before the remote-data stages, install the rest:
pip install -r requirements.txt           # pyarrow, lmfit, netCDF4, scikit-learn, xgboost, joblib
python stages/01_soils/run.py
# ... 02_climate, 03_terrain, 04_phenology, 05_compile, 06_model

# 4) Optional phenology imagery-source comparison:
python stages/04_phenology/run.py --imagery-source landsat --extract-only
python stages/04_phenology/run.py --imagery-source landsat --fit-only
python stages/05_compile/run.py --imagery-source landsat
python stages/06_model/run.py --imagery-source landsat --views phenology,hybrid
python stages/07_testing/run.py --imagery-source landsat --view phenology

python stages/04_phenology/run.py --imagery-source sentinel --extract-only
python stages/04_phenology/run.py --imagery-source sentinel --fit-only
python stages/05_compile/run.py --imagery-source sentinel
python stages/06_model/run.py --imagery-source sentinel --views phenology,hybrid
python stages/07_testing/run.py --imagery-source sentinel --view phenology

python stages/11_imagery_comparison/run.py

Design decisions (see CONTRADICTIONS.md for the full list)

  • —Per-observation-year imagery — each plot uses HLS from its own survey year (2023/24/25).
  • —USGS 2025 is an external held-out test — never trained on; scored separately as the honest benchmark, alongside the pooled random test + spatial CV on the other sources.
  • —Presence = cover ≥ 20% (heavy-vs-light), trace-VEDU kept as negatives.
  • —Reproducibility: rerun any stage → cache hits in seconds unless an input, a param, or a step's CODE_VERSION changed. python -m vedu.manifest status shows what is current vs stale.
  • —Ming data is currently enabled: Stage 00 includes Ming Crow points plus Extra BR point/polygon representatives. Crow polygons stay disabled to avoid double-counting paired Crow observations. To return to the old primary-only dataset, set the three enabled Ming source entries to enabled: false in config/sources.yaml.

Status

Foundation complete (library + config + docs). Stages build one at a time — see `docs/HANDOFF.md` build-status checklist.