Team Ai
Apppublic

nates788/gpt-335M

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes
App README

The model uses a decoder-only GPT-2-style architecture, with multi-head self-attention and MLP blocks stacked across ~22 layers, 16 attention heads, and a 1024-dimensional embedding size, operating over a ~50k vocabulary and a context window of 1k tokens. It was trained autoregressively with a standard next-token prediction objective, using FineWeb-Edu as the primary training corpus. Fine-tuning was performed via LoRA adapters applied to attention and MLP projection layers.

Trained on the FineWeb-Edu dataset, it excels at producing textbook-style descriptions, high-level conceptual overviews, and tutorial-like explanations across domains such as computer science, mathematics, natural sciences, and general academic topics. Rather than being optimized for strict instruction following, exact arithmetic, or conversational alignment, the model is best understood as an educational text continuation model: it generates coherent, informative prose that mirrors how concepts are explained in learning materials. As a result, it is well-suited for explaining ideas and summarizing topics, while being less reliable for logic and arithmetic.