Team Ai
Modelpublic

kihyounghan/workshop-pretraining-v2

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes1.3kdownloads
Model Card

workshop-pretraining-v2

A small (123.6M parameters) decoder-only Transformer language model, pretrained from scratch for only 100 optimizer steps as part of a hands-on LLM pretraining workshop.

This is an educational checkpoint, not a usable language model. It was trained on about 52 million tokens (roughly 0.5% of the 10B-token plan in the workshop notebook), so it produces mostly frequent words with no coherent meaning. See Limitations.

Model details

ItemValue
ArchitectureDecoder-only Transformer (GPT-2 sized, with modern components)
Parameters123,587,328
Layers / hidden size / heads12 / 768 / 12 (head dim 64)
Context length1024 tokens
Vocabulary50,304 (GPT-2 BPE with 50,257 tokens, padded to a multiple of 64)
Feed-forward4x expansion, ReLU² activation, no biases
NormalizationRMSNorm, pre-norm
Positional encodingRotary embeddings (RoPE, theta = 10000)
Output headTied with the input embedding, tanh logit soft-capping (cap = 30)
Dropout0.1 (training only)
Tokenizertiktoken GPT-2 encoding (`<endoftext>` = 50256)

This is a custom architecture, not GPT2LMHeadModel. Loading it requires trust_remote_code=True. The model code is in configuration_gpt2workshop.py and modeling_gpt2workshop.py in this repository.

Training

ItemValue
DataHuggingFaceFW/fineweb-edu, sample-10BT subset, streamed and tokenized into uint16 shards
Training length100 optimizer steps
Tokens per step524,288 (micro-batch 4 x 1024 tokens x 128 gradient-accumulation steps)
Tokens seenabout 52.4M
OptimizerAdamW, betas (0.9, 0.95), weight decay 0.1 (not applied to norms and 1-D parameters), fused
Learning ratePeak 6e-4, linear warmup for 10 steps, then cosine decay to 6e-5
Gradient clipping1.0
Precisionbfloat16 autocast
Hardware1x NVIDIA GeForce RTX 3080 Laptop GPU (8 GB), about 93 minutes, about 9,000 tokens/s
FrameworkPyTorch (no torch.compile)

Results

MetricValue
Training loss, step 108.58
Training loss, step 507.28
Training loss, step 1006.97
LAMBADA accuracy (steps 50 and 100)0.0000

For reference, a uniform random guess over this vocabulary gives a loss of about 10.8.

How to use

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "kihyounghan/workshop-pretraining-v2"
device = "cuda" if torch.cuda.is_available() else "cpu"

tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).to(device)
model.eval()

prompt = "The best way to learn programming is"
input_ids = tokenizer.encode(prompt, return_tensors="pt").to(device)

with torch.no_grad():
    output_ids = model.generate(
        input_ids, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=50
    )

print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

trust_remote_code=True runs the Python files stored in this repository. Review them first, or pin a specific commit with revision="<commit hash>".

Example output

Prompt: The theory of relativity states that

The theory of relativity states that you, they the research and the high of the following of ...

The model has learned which words are frequent in English, but not grammar or meaning yet.

Limitations

  • —Trained for 100 steps only, so outputs are repetitive, ungrammatical and unrelated to the prompt.
  • —Not evaluated for safety, factuality or bias. It is not suitable for any real application.
  • —Trained on English web text only. The training data may contain biases and errors.
  • —Intended for learning and for testing the pretraining, export and upload pipeline.

한국어 요약

LLM 사전학습(pretraining) 워크숍 실습용으로 처음부터 학습한 약 1.24억 파라미터 모델입니다. 100 스텝(약 5,200만 토큰)만 학습해서 실제로 쓸 수 있는 수준이 아니며, 학습 → 저장 → Hugging Face 업로드 과정을 확인하기 위한 예제 체크포인트입니다.