reaperdoesntknow/gemma-270m-math-reasoner
04.8k
gemma-270m-math-reasoner
Experimental checkpoint. A 270M-parameter Gemma 3 fine-tuned to produce step-by-step math reasoning. It was trained as a quick capacity test and published mainly as a checkpoint — treat it as a research artifact, not a production math model.
What to expect
- It has learned the form of reasoning: it writes out steps and works toward a final answer.
- At 270M parameters it often can't carry the arithmetic through. Expect confident-looking chains that go wrong mid-way, especially on multi-step word problems.
- Useful for: studying how far reasoning-format training transfers at very small scale, edge/on-device experiments, and as a baseline against larger distills.
A larger (1B) variant is planned.
Quick start
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "reaperdoesntknow/gemma-270m-math-reasoner"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
messages = [{"role": "user", "content": "A shop sells pens at 3 for $2. How much do 12 pens cost?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
out = model.generate(inputs, max_new_tokens=256)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))Training
- Base: google/gemma-3-270m
- Data: <!-- fill in: dataset(s) / trace source -->
- Method: <!-- fill in: SFT / distillation, epochs, lr -->
- Hardware: <!-- fill in -->
Evaluation
<!-- fill in after a quick GSM8K subset run, e.g. "GSM8K (200-problem sample): X%" --> Not yet formally evaluated.
More from this author
Published by Convergent Intelligence LLC. <!-- cix-keeper-ts:2026-10-03T13:17:08Z -->
