kihyounghan/workshop-pretraining-v2
workshop-pretraining-v2
A small (123.6M parameters) decoder-only Transformer language model, pretrained from scratch for only 100 optimizer steps as part of a hands-on LLM pretraining workshop.
This is an educational checkpoint, not a usable language model. It was trained on about 52 million tokens (roughly 0.5% of the 10B-token plan in the workshop notebook), so it produces mostly frequent words with no coherent meaning. See Limitations.
Model details
This is a custom architecture, not GPT2LMHeadModel. Loading it requires trust_remote_code=True. The model code is in configuration_gpt2workshop.py and modeling_gpt2workshop.py in this repository.
Training
Results
For reference, a uniform random guess over this vocabulary gives a loss of about 10.8.
How to use
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "kihyounghan/workshop-pretraining-v2"
device = "cuda" if torch.cuda.is_available() else "cpu"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).to(device)
model.eval()
prompt = "The best way to learn programming is"
input_ids = tokenizer.encode(prompt, return_tensors="pt").to(device)
with torch.no_grad():
output_ids = model.generate(
input_ids, max_new_tokens=50, do_sample=True, temperature=0.8, top_k=50
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))trust_remote_code=True runs the Python files stored in this repository. Review them first, or pin a specific commit with revision="<commit hash>".
Example output
Prompt: The theory of relativity states that
The theory of relativity states that you, they the research and the high of the following of ...The model has learned which words are frequent in English, but not grammar or meaning yet.
Limitations
- Trained for 100 steps only, so outputs are repetitive, ungrammatical and unrelated to the prompt.
- Not evaluated for safety, factuality or bias. It is not suitable for any real application.
- Trained on English web text only. The training data may contain biases and errors.
- Intended for learning and for testing the pretraining, export and upload pipeline.
한국어 요약
LLM 사전학습(pretraining) 워크숍 실습용으로 처음부터 학습한 약 1.24억 파라미터 모델입니다. 100 스텝(약 5,200만 토큰)만 학습해서 실제로 쓸 수 있는 수준이 아니며, 학습 → 저장 → Hugging Face 업로드 과정을 확인하기 위한 예제 체크포인트입니다.
