Team Ai
Modelpublic

tsilva/Level2-3_stable-baselines3-ppo_ff330a59

sourceHugging Facemitupdated 9d agoView on Hugging Face
0likes14downloads
Model Card

SuperMarioBros-Nes-v0 — Level2-3 — PPO

Stable-Baselines3 PPO policy for SuperMarioBros-Nes-v0 Level2-3, trained and evaluated with `rlab`.

At a Glance

ItemValue
TaskComplete SuperMarioBros-Nes-v0 Level2-3
Providersupermariobrosnes-turbo
Algorithmppo
CheckpointStep 4500000
Evaluationstochastic full evaluation, 100 episodes
Successminimum 98.0%, mean 98.0%
Mean return3516.013
Releasev1
PreviewRoot replay.mp4
YouTubeWatch on YouTube

Quick Start

bash
git clone https://github.com/tsilva/rlab
cd rlab
git checkout 1e18630ff20f40c4038d6bee5b023ff704c7277b
uv sync --frozen

Import the ROM, then play or evaluate the immutable checkpoint:

bash
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level2-3_stable-baselines3-ppo_ff330a59/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level2-3_stable-baselines3-ppo_ff330a59/resolve/v1/model.zip

Evaluation

Action selection was stochastic under the published evaluation environment contract.

StartEpisodesSuccessesSuccess rateMean return
Level2-31009898.0%3516.013

Environment and Policy Contract

ItemValue
Environmentsupermariobrosnes-turbo:SuperMarioBros-Nes-v0
Environment hashsha256:1c0bc76e09b9e82e9505321e5592e136154236356c081c488c1bec49057f7d5a
Preprocessing{"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0}
Action contract{"set":"simple"}

Provenance

ItemValue
Sourcerlab
RunLevel2-3_base_s1_20260704T121248Z
Recipebase
Seed1
Source commit1e18630ff20f40c4038d6bee5b023ff704c7277b
Evaluated artifacttsilva/SuperMarioBros-Nes-v0/Level2-3_base_s1_20260704T121248Z-checkpoint:latest

Files

FilePurpose
model.zipStable-Baselines3 policy checkpoint
model.jsonVersioned checkpoint identity, policy type, provenance, and recipe binding
recipe.jsonVersioned execution and evaluation contract
release_manifest.jsonRelease identity, evaluation evidence, and artifact hashes
replay.mp4Browser-safe representative episode
LICENSELicense for rlab-authored policy weights and publication material

Limitations

Evaluation establishes performance only for the published environment hash, start distribution, policy preprocessing, and action-selection protocol. It does not establish generalization to other levels, environments, ROM revisions, or contracts.

Licensing

The rlab-authored policy weights and publication material are licensed under the MIT License in LICENSE. Emulator/runtime software and game assets remain governed by their own licenses and terms. This repository does not redistribute a game ROM.

Policy Lineage

This is a legacy rlab policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.

  • —Trainer: Stable-Baselines3
  • —Algorithm: PPO
  • —Model class: stable_baselines3.ppo.ppo.PPO
  • —Full lineage digest: ff330a5907addffb33908b093445dbb9c23c01b7dd6e3f9dd95159fe73827d67
  • —Immutable release: hf://tsilva/Level2-3_stable-baselines3-ppo_ff330a59@v1
  • —Exact checkpoint tag: checkpoint-4500000

The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.