tsilva/Level2-3_stable-baselines3-ppo_ff330a59
SuperMarioBros-Nes-v0 — Level2-3 — PPO
Stable-Baselines3 PPO policy for SuperMarioBros-Nes-v0 Level2-3, trained and evaluated with `rlab`.
At a Glance
Quick Start
git clone https://github.com/tsilva/rlab
cd rlab
git checkout 1e18630ff20f40c4038d6bee5b023ff704c7277b
uv sync --frozenImport the ROM, then play or evaluate the immutable checkpoint:
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level2-3_stable-baselines3-ppo_ff330a59/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level2-3_stable-baselines3-ppo_ff330a59/resolve/v1/model.zipEvaluation
Action selection was stochastic under the published evaluation environment contract.
Environment and Policy Contract
Provenance
Files
Limitations
Evaluation establishes performance only for the published environment hash, start distribution, policy preprocessing, and action-selection protocol. It does not establish generalization to other levels, environments, ROM revisions, or contracts.
Licensing
The rlab-authored policy weights and publication material are licensed under the MIT License in LICENSE. Emulator/runtime software and game assets remain governed by their own licenses and terms. This repository does not redistribute a game ROM.
Policy Lineage
This is a legacy rlab policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.
- Trainer:
Stable-Baselines3 - Algorithm:
PPO - Model class:
stable_baselines3.ppo.ppo.PPO - Full lineage digest:
ff330a5907addffb33908b093445dbb9c23c01b7dd6e3f9dd95159fe73827d67 - Immutable release:
hf://tsilva/Level2-3_stable-baselines3-ppo_ff330a59@v1 - Exact checkpoint tag:
checkpoint-4500000
The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.
