FluidInference/intern-decision-0.8b-showdown-coreml
Intern-Decision-0.8B, fine-tuned for Pokémon Showdown, for Core ML
internlm/Intern-Decision-0.8B (Shanghai AI Laboratory, Apache-2.0) fine-tuned to choose Pokémon Showdown battle actions, exported for Core ML. The stock model plays at chance; this one learned from its own 4B sibling, internlm/Intern-Decision-4B, over one night on a MacBook Pro. Inspired by SGLang's Qwen3.8-27B FireRed run; this is the small, local end of that idea, on the Showdown simulator rather than the game.
How it was trained
- Teacher. Intern-Decision-4B played 220
gen9randombattlebattles on a local Showdown server through poke-env against poke-env's random, max-base-power and simple-heuristics players and itself; every turn's request and probability vector was logged (6,974 decisions). - Request. A compact JSON state (both teams, HP, status, boosts, field) and one
choicequestion whose options are the legal moves and switches with their facts (type, category, power, STAB, effectiveness, accuracy, PP; switch matchups). Exactly Intern-Decision's own wire format, rendered by the checkpoint'sinference.py. - Student. LoRA r=32 on every language-model linear layer, 1.5 epochs, KL(teacher ‖ student) on the temperature-scaled restricted softmax at the
<decision>marker, option order shuffled per sample. 200 minutes on an M5 Pro (PyTorch MPS). Merged into the weights before export. Held-out agreement with the 4B's top choice: 35% → 80%. - Not used. The 27B, any GPU server, or any game ROM.
Results (gen9 random battles, this model's side first)
It matches its teacher and loses to poke-env's heuristic bot most of the time, like the teacher does. Latency on an M5 Pro GPU: 89 ms per decision in the 512-token bucket, 125 ms in the 640, 180 ms in the 1024 (typical battle requests are 430–560 tokens).
Files
int8 (per-channel, weight-only) is the recommended download: measured on the Swift runtime with one bucket in use it is about 0.85 GB in memory (374 MB process footprint + 480 MB mapped weights) against about 1.5 GB for fp16, at the same 90 ms per decision, and it played 15-0 against poke-env's max-base-power player where fp16 played 14-1 (15 battles each). FluidUse's .showdown snapshot uses the int8 buckets.
Same inputs and outputs as FluidInference/intern-decision-0.8b-coreml: hidden [1, L, 1024], cos / sin [L, 64], field_onehot [F, L] → logits [F, 62] over the answer symbols; softmax over the first n symbols, log-probabilities divided by the temperature. It is a general typed-decision model that happens to know Pokémon: any state and question set works, but the fine-tune was only evaluated on battles. Runtime: InternDecisionManager in FluidUse; harness, training and export code in mobius under models/computer-use/intern-decision-0.8b/showdown/.
