Team Ai
Modelpublic

reaperdoesntknow/gemma-270m-math-reasoner

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes4.8kdownloads
Model Card

gemma-270m-math-reasoner

Experimental checkpoint. A 270M-parameter Gemma 3 fine-tuned to produce step-by-step math reasoning. It was trained as a quick capacity test and published mainly as a checkpoint — treat it as a research artifact, not a production math model.

What to expect

  • —It has learned the form of reasoning: it writes out steps and works toward a final answer.
  • —At 270M parameters it often can't carry the arithmetic through. Expect confident-looking chains that go wrong mid-way, especially on multi-step word problems.
  • —Useful for: studying how far reasoning-format training transfers at very small scale, edge/on-device experiments, and as a baseline against larger distills.

A larger (1B) variant is planned.

Quick start

python
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "reaperdoesntknow/gemma-270m-math-reasoner"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

messages = [{"role": "user", "content": "A shop sells pens at 3 for $2. How much do 12 pens cost?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
out = model.generate(inputs, max_new_tokens=256)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

Training

  • —Base: google/gemma-3-270m
  • —Data: <!-- fill in: dataset(s) / trace source -->
  • —Method: <!-- fill in: SFT / distillation, epochs, lr -->
  • —Hardware: <!-- fill in -->

Evaluation

<!-- fill in after a quick GSM8K subset run, e.g. "GSM8K (200-problem sample): X%" --> Not yet formally evaluated.

More from this author

Published by Convergent Intelligence LLC. <!-- cix-keeper-ts:2026-10-03T13:17:08Z -->