Team Ai
Modelpublic

FluidInference/intern-decision-0.8b-showdown-coreml

sourceHugging Faceapache-2.0updated 6d agoView on Hugging Face
0likes65downloads
Model Card

Intern-Decision-0.8B, fine-tuned for Pokémon Showdown, for Core ML

internlm/Intern-Decision-0.8B (Shanghai AI Laboratory, Apache-2.0) fine-tuned to choose Pokémon Showdown battle actions, exported for Core ML. The stock model plays at chance; this one learned from its own 4B sibling, internlm/Intern-Decision-4B, over one night on a MacBook Pro. Inspired by SGLang's Qwen3.8-27B FireRed run; this is the small, local end of that idea, on the Showdown simulator rather than the game.

How it was trained

  • —Teacher. Intern-Decision-4B played 220 gen9randombattle battles on a local Showdown server through poke-env against poke-env's random, max-base-power and simple-heuristics players and itself; every turn's request and probability vector was logged (6,974 decisions).
  • —Request. A compact JSON state (both teams, HP, status, boosts, field) and one choice question whose options are the legal moves and switches with their facts (type, category, power, STAB, effectiveness, accuracy, PP; switch matchups). Exactly Intern-Decision's own wire format, rendered by the checkpoint's inference.py.
  • —Student. LoRA r=32 on every language-model linear layer, 1.5 epochs, KL(teacher ‖ student) on the temperature-scaled restricted softmax at the <decision> marker, option order shuffled per sample. 200 minutes on an M5 Pro (PyTorch MPS). Merged into the weights before export. Held-out agreement with the 4B's top choice: 35% → 80%.
  • —Not used. The 27B, any GPU server, or any game ROM.

Results (gen9 random battles, this model's side first)

Playervs randomvs max-base-powervs simple heuristics
stock Intern-Decision-0.8B (Core ML), 10 battles2-31-91-9
Intern-Decision-4B teacher (PyTorch), 60 battles58-245-1516-44
this model (Core ML), 30 battles9-1 (10 played)24-69-21

It matches its teacher and loses to poke-env's heuristic bot most of the time, like the teacher does. Latency on an M5 Pro GPU: 89 ms per decision in the 512-token bucket, 125 ms in the 640, 180 ms in the 1024 (typical battle requests are 430–560 tokens).

Files

PathWhat
L512_F8/, L640_F8/, L1024_F16/DecisionRow_w8.mlpackage (int8 weights, 480 MB) and DecisionRow_fp16.mlpackage (955 MB) + config.json per bucket (tokens × fields)
multi/DecisionRow_w8.mlpackageall three buckets as one weight-shared multifunction package (488 MB, functions L512_F8 / L640_F8 / L1024_F16; needs macOS 15 / iOS 18 to select a function)
embeddings.f16token embeddings (fp16, 248,320 × 1,024), gathered on the host
tokenizer.jsonthe checkpoint's tokenizer (Qwen3.5 + <decision>)

int8 (per-channel, weight-only) is the recommended download: measured on the Swift runtime with one bucket in use it is about 0.85 GB in memory (374 MB process footprint + 480 MB mapped weights) against about 1.5 GB for fp16, at the same 90 ms per decision, and it played 15-0 against poke-env's max-base-power player where fp16 played 14-1 (15 battles each). FluidUse's .showdown snapshot uses the int8 buckets.

Same inputs and outputs as FluidInference/intern-decision-0.8b-coreml: hidden [1, L, 1024], cos / sin [L, 64], field_onehot [F, L] → logits [F, 62] over the answer symbols; softmax over the first n symbols, log-probabilities divided by the temperature. It is a general typed-decision model that happens to know Pokémon: any state and question set works, but the fine-tune was only evaluated on battles. Runtime: InternDecisionManager in FluidUse; harness, training and export code in mobius under models/computer-use/intern-decision-0.8b/showdown/.