Team Ai
Datasetpublic

Dhevenddra/slm-lab-data

slm-lab-data Everything slm-lab produced that is not a model: the synthetic task it generated, the corpora it packed, the tokenizers it trained from scratch, and every published result file. What it is for. Two things a reader can actually do with it. The browsable configs below are training data with ground truth that is correct by construction — the expense→JSON task is generated by scripts/gen_json_task.py, so every target is exact, including the computed dates. And results/… See the full description on the dataset page: https://huggingface.co/datasets/Dhevenddra/slm-lab-data.

sourceHugging Facemitupdated 16h agoView on Hugging Face
1likes646downloads
Dataset Card

slm-lab-data

Everything slm-lab produced that is not a model: the synthetic task it generated, the corpora it packed, the tokenizers it trained from scratch, and every published result file.

What it is for. Two things a reader can actually do with it. The browsable configs below are training data with ground truth that is correct by construction — the expense→JSON task is generated by scripts/gen_json_task.py, so every target is exact, including the computed dates. And results/ holds every measurement this project made, including the negative ones, in the form they were published.

Browse it

Six configs are set up for the dataset viewer. expense_to_json is the default and the place to start.

configrows arewhat it was used for
`expense_to_json`a messy expense line, a reference date, the exact JSON target, and the chat messages built from themthe fine-tuning task: train / validation / test
expense_to_json_preference_promptsthe same shape, on held-out linesthe prompts the preference pairs were generated from
preference_pairsprompt, chosen, rejected, plus which fields the rejected one got wrongDPO training
chat_sftchat messagesinstruction tuning the 12.6M model
chat_sft_v1chat messagesthe superseded first attempt, kept because the pair of runs is the result
scored_generationsmodel, prompt, generated text, hand-scored responsivenessthe evaluation-by-reading pass

The rest of the repo is files rather than tables, and the viewer is told to leave them alone:

pathwhat it is
tinystories_<v>.jsonbyte-level BPE tokenizers trained from scratch at vocab 2048 / 8192 / 32768
tinystories_<v>_{train,val}.binTinyStories packed to uint16 token streams
chat_8192*the chat-formatted corpus for mid-training, packed
sweeps/one metrics.json per ablation arm
results/every published write-up, table, figure and score file

The results are the point

results/ holds the project's measured findings in full, including the ones that went the wrong way:

  • —summary.md — 24 runs, 7 ablations, judged against a measured seed noise floor of 0.0146 nats. Two of seven showed no detectable difference and are reported as such.
  • —ablation_vocab_size.md — and the note explaining that correcting the bits-per-byte metric reversed this ablation's verdict, because the original was per-token loss times a shared constant and so preserved the exact artifact it existed to remove.
  • —phase08_compare.md — every production quantization method on one table, including the KL column almost nobody publishes.
  • —phase08_granularity.md — quantization granularity measured by hand on one real weight matrix, with the outlier experiment.
  • —phase09_scaling.md — the capability floor, bracketed across a 98× parameter range, with the chance baseline next to every score.

A note on the tokenizers

Three vocabularies were trained so the choice could be measured rather than assumed. The 32768 arm wins on bits-per-byte — but its embedding table is 12.6M parameters, which at this project's budget is the entire model, leaving 17K parameters for the layers. The measured winner and the usable choice are not the same thing, and the results say so.

Limitations

Specific, and written out in full — this is the section a knowledgeable reader checks first.

  • —Everything here is synthetic. TinyStories, a public chat corpus, and a task generated by this project's own script. There is no personal data in any of it, and no human-written text that was not already public.
  • —The task is narrow by design. One schema, seven fields, one domain. It was chosen because its ground truth can be known exactly, not because it is representative.
  • —The generated corpus inherits its generator's blind spots. The merchant names, phrasings and date expressions come from fixed vocabularies in gen_json_task.py, so a model that fits this task has not been shown the variety real expense text has.
  • —Models trained on it are research artifacts for learning, not products.

Where this came from

Part of slm-lab — a language model built end to end on free compute: a hand-written BPE tokenizer, a transformer from scratch, a controlled ablation study with a measured noise floor, post-training, QLoRA on a 1B model, DPO, and quantization written by hand and benchmarked against every production method.

  • —Start here: the SLM Lab collection lists every repo in the order it was built, with a one-line finding on each.
  • —Code, guide and the full failure log: https://github.com/Dhevenddra/slm-lab
  • —Every result file, dataset and tokenizer: <https://huggingface.co/datasets/Dhevenddra/slm-lab-data>

This repo is the project's data, tokenizers and published results stage; the published write-up for it is under results/ in the dataset repo, named for that phase.

The project's own failure log (FAILURES.md, 26 entries) is the part most worth reading. The recurring shape: the instrument measured something correlated with the question rather than equal to it.