SLM-Archive/NanoDex-1M
<div align="center"> <h3 align="center">Model archived by:</h3>
<table align="center"> <tr> <td align="center" valign="middle" width="140"> <a href="https://huggingface.co/DedeProGames" target="_blank"> <img src="https://cdn-avatars.huggingface.co/v1/production/uploads/685ea8ff7b4139b6845ce395/Im--QSnbrnAhHPPhpX8L0.png" width="80" height="80" alt="DedeProGames" style="display:block; margin:0 auto; border-radius:50%; object-fit:cover;" /> <br /> <b>DedeProGames</b> </a> </td> </tr> </table> </div>
NanoDex-1M
A 1,062,272-parameter decoder-only language model pre-trained from scratch on fineweb-edu, using the NanoDex Trainer Space.
Architecture
A standard LlamaForCausalLM decoder-only transformer — SiLU MLP, RMSNorm, rotary position embeddings, grouped-query attention, tied embeddings, no biases — scaled down in width and depth to fit the parameter budget.
Training
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("DedeProGames/NanoDex-1M")
model = AutoModelForCausalLM.from_pretrained("DedeProGames/NanoDex-1M")
ids = tok("The mitochondria is", return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, max_new_tokens=60, do_sample=True,
temperature=0.8, top_k=50)[0]))Caveats
This is a nano-scale research artifact. At this parameter count and token budget the model learns word shapes, common collocations and a little syntax — it is not a useful assistant and its output is not factual. It exists to make "pre-train a transformer from scratch" something you can actually watch happen.
