Team Ai
Datasetpublic

EXD-AI/episode-07-tokenization

Ep. 7 — Tokenization & Embeddings How does a sentence become a list of vectors — and what does that space look like? Contents Notebook: tokenization_and_embeddings.ipynb Sections # Topic 0 Full pipeline — text → tokens → atom (live vLLM inference) 1 BPE tokenization from scratch 2 The embedding matrix — loading Qwen's actual weights 3 Exploring embedding space (t-SNE, cosine similarity, nearest neighbors) Hardware… See the full description on the dataset page: https://huggingface.co/datasets/EXD-AI/episode-07-tokenization.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes10downloads
Dataset Card

Ep. 7 — Tokenization & Embeddings

How does a sentence become a list of vectors — and what does that space look like?

Contents

  • —Notebook: tokenization_and_embeddings.ipynb

Sections

#Topic
0Full pipeline — text → tokens → atom (live vLLM inference)
1BPE tokenization from scratch
2The embedding matrix — loading Qwen's actual weights
3Exploring embedding space (t-SNE, cosine similarity, nearest neighbors)

Hardware

The notebook connects to the live vLLM server on atom (DGX Spark, Grace-Blackwell) for the inference section. Tokenizer and embeddings are loaded directly from Qwen/Qwen3.6-35B-A3B weights.

Related