BananaMind/normal-model-ablation
0357
Standard attention ablation model (1.50M params, 3072 context)
Part of a controlled ablation: BananaMind/dsa-model-ablation (DSA) vs BananaMind/normal-model-ablation (standard attention). Same seed, data order, tokenizer, architecture and hyper-parameters. Only the attention differs.
- Attention: Standard dense causal multi-head attention.
- Architecture: 5 layers, hidden 128, 4 heads, SwiGLU 336, RoPE, RMSNorm, tied embeddings, vocab 4096 (custom BPE)
- Context length: 3072
- Params: 1,498,496
- Data: HuggingFaceFW/fineweb-edu (sample-10BT), 2000M tokens, 1 epoch, seed 1337
- Optimiser: AdamW lr 0.002 (cosine), batch 16x3072 tokens, bf16 autocast
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "BananaMind/normal-model-ablation"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)(No KV cache is implemented; generate recomputes the full sequence each step.)
Benchmarks (0-shot, lm-evaluation-harness; acc / acc_norm where available)
Per-model raw results: benchmarks/results.json. Comparison with the other model: benchmarks/comparison.md (other repo: BananaMind/dsa-model-ablation). Training curve: train_log.json.
Note: at this size models are close to chance on these benchmarks, and the benchmark prompts are shorter than the top-k (512), where DSA selects every token and behaves identically to dense attention at inference, so differences mostly reflect how training differed.
