Team Ai
Datasetpublic

alexkstern/github-code-nanochatbpe-1B

github-code-nanochatbpe-1B GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 1,000,000,000 val.bin val 10,000,000 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes63downloads
Dataset Card

github-code-nanochatbpe-1B

GitHub Code (all-all) (from `codeparrot/github-code`), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training.

filesplittokens
train.bintrain1,000,000,000
val.binval10,000,000

train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry the full metadata. The tokenizer/ files (when present) are the exact tokenizer used to produce these ids.

Load a bin with the standard Hugging Face downloader:

python
from huggingface_hub import hf_hub_download
import numpy as np

path = hf_hub_download(repo_id="alexkstern/github-code-nanochatbpe-1B", filename="train.bin", repo_type="dataset")
tokens = np.memmap(path, dtype="uint16", mode="r")