Dhevenddra/slm-lab-data
slm-lab-data Everything slm-lab produced that is not a model: the synthetic task it generated, the corpora it packed, the tokenizers it trained from scratch, and every published result file. What it is for. Two things a reader can actually do with it. The browsable configs below are training data with ground truth that is correct by construction — the expense→JSON task is generated by scripts/gen_json_task.py, so every target is exact, including the computed dates. And results/… See the full description on the dataset page: https://huggingface.co/datasets/Dhevenddra/slm-lab-data.
slm-lab-data
Everything slm-lab produced that is not a model: the synthetic task it generated, the corpora it packed, the tokenizers it trained from scratch, and every published result file.
What it is for. Two things a reader can actually do with it. The browsable configs below are training data with ground truth that is correct by construction — the expense→JSON task is generated by scripts/gen_json_task.py, so every target is exact, including the computed dates. And results/ holds every measurement this project made, including the negative ones, in the form they were published.
Browse it
Six configs are set up for the dataset viewer. expense_to_json is the default and the place to start.
The rest of the repo is files rather than tables, and the viewer is told to leave them alone:
The results are the point
results/ holds the project's measured findings in full, including the ones that went the wrong way:
summary.md— 24 runs, 7 ablations, judged against a measured seed noise floor of 0.0146 nats. Two of seven showed no detectable difference and are reported as such.ablation_vocab_size.md— and the note explaining that correcting the bits-per-byte metric reversed this ablation's verdict, because the original was per-token loss times a shared constant and so preserved the exact artifact it existed to remove.phase08_compare.md— every production quantization method on one table, including the KL column almost nobody publishes.phase08_granularity.md— quantization granularity measured by hand on one real weight matrix, with the outlier experiment.phase09_scaling.md— the capability floor, bracketed across a 98× parameter range, with the chance baseline next to every score.
A note on the tokenizers
Three vocabularies were trained so the choice could be measured rather than assumed. The 32768 arm wins on bits-per-byte — but its embedding table is 12.6M parameters, which at this project's budget is the entire model, leaving 17K parameters for the layers. The measured winner and the usable choice are not the same thing, and the results say so.
Limitations
Specific, and written out in full — this is the section a knowledgeable reader checks first.
- Everything here is synthetic. TinyStories, a public chat corpus, and a task generated by this project's own script. There is no personal data in any of it, and no human-written text that was not already public.
- The task is narrow by design. One schema, seven fields, one domain. It was chosen because its ground truth can be known exactly, not because it is representative.
- The generated corpus inherits its generator's blind spots. The merchant names, phrasings and date expressions come from fixed vocabularies in
gen_json_task.py, so a model that fits this task has not been shown the variety real expense text has. - Models trained on it are research artifacts for learning, not products.
Where this came from
Part of slm-lab — a language model built end to end on free compute: a hand-written BPE tokenizer, a transformer from scratch, a controlled ablation study with a measured noise floor, post-training, QLoRA on a 1B model, DPO, and quantization written by hand and benchmarked against every production method.
- Start here: the SLM Lab collection lists every repo in the order it was built, with a one-line finding on each.
- Code, guide and the full failure log: https://github.com/Dhevenddra/slm-lab
- Every result file, dataset and tokenizer: <https://huggingface.co/datasets/Dhevenddra/slm-lab-data>
This repo is the project's data, tokenizers and published results stage; the published write-up for it is under results/ in the dataset repo, named for that phase.
The project's own failure log (FAILURES.md, 26 entries) is the part most worth reading. The recurring shape: the instrument measured something correlated with the question rather than equal to it.
