Team Ai
Datasetpublic

EXDai/episode-07-tokenization

Ep. 7 — Tokenization & Embeddings How does a sentence become a list of vectors — and what does that space look like? Contents Notebook: tokenization_and_embeddings.ipynb Sections # Topic 0 Full pipeline — text → tokens → atom (live vLLM inference) 1 BPE tokenization from scratch 2 The embedding matrix — loading Qwen's actual weights 3 Exploring embedding space (t-SNE, cosine similarity, nearest neighbors) Related… See the full description on the dataset page: https://huggingface.co/datasets/EXDai/episode-07-tokenization.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes3downloads
Dataset Card

Ep. 7 — Tokenization & Embeddings

How does a sentence become a list of vectors — and what does that space look like?

Contents

  • —Notebook: tokenization_and_embeddings.ipynb

Sections

#Topic
0Full pipeline — text → tokens → atom (live vLLM inference)
1BPE tokenization from scratch
2The embedding matrix — loading Qwen's actual weights
3Exploring embedding space (t-SNE, cosine similarity, nearest neighbors)

Related