Team Ai
Modelpublic

tsilva/Level3-4_stable-baselines3-ppo_4141caab

sourceHugging Facemitupdated 9d agoView on Hugging Face
0likes10downloads
Model Card

SuperMarioBros-Nes-v0 — Level3-4 — PPO

Stable-Baselines3 PPO policy for SuperMarioBros-Nes-v0 Level3-4, trained and evaluated with `rlab`.

At a Glance

ItemValue
TaskComplete SuperMarioBros-Nes-v0 Level3-4
Providersupermariobrosnes-turbo
Algorithmppo
CheckpointStep 3220272
Evaluationstochastic full evaluation, 100 episodes
Successminimum 96.0%, mean 96.0%
Mean return2283.246
Releasev1
PreviewRoot replay.mp4
YouTubeWatch on YouTube

Quick Start

bash
git clone https://github.com/tsilva/rlab
cd rlab
git checkout a21e8fc154ecf3e47e39f1fc523398b575cd12ed
uv sync --frozen

Import the ROM, then play or evaluate the immutable checkpoint:

bash
uv run rlab import-roms ~/roms --game SuperMarioBros-Nes-v0
uv run rlab play https://huggingface.co/tsilva/Level3-4_stable-baselines3-ppo_4141caab/resolve/v1/model.zip
uv run rlab eval https://huggingface.co/tsilva/Level3-4_stable-baselines3-ppo_4141caab/resolve/v1/model.zip

Evaluation

Action selection was stochastic under the published evaluation environment contract.

StartEpisodesSuccessesSuccess rateMean return
Level3-41009696.0%2283.246

Environment and Policy Contract

ItemValue
Environmentsupermariobrosnes-turbo:SuperMarioBros-Nes-v0
Environment hashsha256:7969db636994526e3c3c1614d378f6de882854a51d54ac9d3ac664fe00b4248a
Preprocessing{"frame_skip":4,"frame_stack":4,"max_pool_frames":false,"obs_copy":"safe_view","obs_crop":[32,0,0,0],"obs_crop_fill":0,"obs_crop_mode":"remove","obs_grayscale":true,"obs_resize":[84,84],"obs_resize_algorithm":"area","pipeline":"supermariobrosnes_turbo_native_vec_env","policy_observation_layout":"channel_first","sticky_action_prob":0.0}
Action contract{"set":"simple"}

Provenance

ItemValue
Sourcerlab
RunLevel3-4_base_s1_20260704T131342Z
Recipebase
Seed1
Source commita21e8fc154ecf3e47e39f1fc523398b575cd12ed
Evaluated artifacttsilva/SuperMarioBros-Nes-v0/Level3-4_base_s1_20260704T131342Z-final:latest

Files

FilePurpose
model.zipStable-Baselines3 policy checkpoint
model.jsonVersioned checkpoint identity, policy type, provenance, and recipe binding
recipe.jsonVersioned execution and evaluation contract
release_manifest.jsonRelease identity, evaluation evidence, and artifact hashes
replay.mp4Browser-safe representative episode
LICENSELicense for rlab-authored policy weights and publication material

Limitations

Evaluation establishes performance only for the published environment hash, start distribution, policy preprocessing, and action-selection protocol. It does not establish generalization to other levels, environments, ROM revisions, or contracts.

Licensing

The rlab-authored policy weights and publication material are licensed under the MIT License in LICENSE. Emulator/runtime software and game assets remain governed by their own licenses and terms. This repository does not redistribute a game ROM.

Policy Lineage

This is a legacy rlab policy trained by Stable-Baselines3 PPO. It is not a GradLab-trained checkpoint and no GradLab compatibility is claimed.

  • —Trainer: Stable-Baselines3
  • —Algorithm: PPO
  • —Model class: stable_baselines3.ppo.ppo.PPO
  • —Full lineage digest: 4141caabd238d66ed39cf0949c208ba5661d8fc16cc20eb1928ee1d24f8fe369
  • —Immutable release: hf://tsilva/Level3-4_stable-baselines3-ppo_4141caab@v1
  • —Exact checkpoint tag: checkpoint-3220272

The mutable main branch contains this corrected card; the original v1 release files, commit, and tag remain unchanged.