Team Ai
Modelpublic

Compactbot/ldt-10m

sourceHugging Faceapache-2.0updated 6d agoView on Hugging Face
0likes1.7kdownloads
Model Card

LDT-10M

A 10M-parameter LLaMA-style language model trained from scratch on 2.6B tokens of FineWeb-Edu.

Architecture

ParameterValue
Params10,284,480
Layers5
d_model320
Heads (Q/KV)5 / 5
Head dim64
FFN (SwiGLU)896
Vocab12,288 (BPE)
Context512
NormRMSNorm
Rope θ10,000
Tied embeddingsyes
PrecisionF32

Training

  • —Data: 2.6B tokens from FineWeb-Edu (HuggingFaceFW/fineweb-edu, sample/10BT, 14 parquet shards)
  • —Steps: 79,375 (batch 64 × seq 512 = 32,768 tok/step)
  • —Optimizer: AdamW (β1=0.9, β2=0.95, wd=0.1)
  • —LR: 3e-5 constant (cosine to 3e-5, effectively flat)
  • —Hardware: RTX 5090, 32 GB
  • —Final val loss: 3.63 (held-out 1.52M-token tail)

Quality

This is a from-scratch training demonstration, not a coherent generator.

Greedy generation produces a grammatical first sentence, then collapses into repetition loops:

Prompt: "The cat sat on the" Output: "The cat sat on the ground, and the other two, and the other two, and the other two, and the other two, and the other two, and the other"

This is expected behavior for a 10M-param model at this training scale. The model has genuinely learned English sentence structure (it beats unigram baseline by a wide margin on held-out perplexity), but sustained coherent generation requires significantly more parameters.

Usage

The model uses a custom architecture (LDTModel) not yet supported by Hugging Face Transformers. To load and run it, use the reference implementation:

python
import torch
from safetensors.torch import load_file
from tokenizers import Tokenizer

# (LDTModel class definition from training script — see CompactAI/ldt-10m-train repo)
model = LDTModel()
model.load_state_dict(load_file("model.safetensors"))
model.eval()

tok = Tokenizer.from_file("tokenizer.json")
ids = tok.encode("The cat sat on the").ids
# ... autoregressive generation ...

Files

FileSizeDescription
model.safetensors41 MBModel weights (F32)
tokenizer.json830 KBBPE tokenizer (12,288 vocab)
config.json287 BArchitecture config