Team Ai
Modelpublic

BananaMind/normal-model-ablation

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes357downloads
Model Card

Standard attention ablation model (1.50M params, 3072 context)

Part of a controlled ablation: BananaMind/dsa-model-ablation (DSA) vs BananaMind/normal-model-ablation (standard attention). Same seed, data order, tokenizer, architecture and hyper-parameters. Only the attention differs.

  • —Attention: Standard dense causal multi-head attention.
  • —Architecture: 5 layers, hidden 128, 4 heads, SwiGLU 336, RoPE, RMSNorm, tied embeddings, vocab 4096 (custom BPE)
  • —Context length: 3072
  • —Params: 1,498,496
  • —Data: HuggingFaceFW/fineweb-edu (sample-10BT), 2000M tokens, 1 epoch, seed 1337
  • —Optimiser: AdamW lr 0.002 (cosine), batch 16x3072 tokens, bf16 autocast

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "BananaMind/normal-model-ablation"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)

(No KV cache is implemented; generate recomputes the full sequence each step.)

Benchmarks (0-shot, lm-evaluation-harness; acc / acc_norm where available)

ModelParamsVal lossVal pplpiqahellaswagarc_easyarc_challenge
DSA1,530,4963.415030.4254.90 / 53.0526.84 / 26.7230.18 / 31.0216.98 / 20.73
Standard1,498,4963.424830.7253.81 / 53.7526.86 / 26.7530.68 / 31.4417.83 / 20.99

Per-model raw results: benchmarks/results.json. Comparison with the other model: benchmarks/comparison.md (other repo: BananaMind/dsa-model-ablation). Training curve: train_log.json.

Note: at this size models are close to chance on these benchmarks, and the benchmark prompts are shorter than the top-k (512), where DSA selects every token and behaves identically to dense attention at inference, so differences mostly reflect how training differed.