xinbenlv/gemma-finetune-webgpu
gemma-finetune-webgpu Built voice/style instruction-tuned datasets used by the gemma-finetune workshop (May 2026, Immersive Commons). Each row is dolly-15k–shaped: {"instruction": "...", "context": "...", "response": "...", "category": "..."} Files file rows upstream recipe shakespeare_15k.jsonl 15,000 HF benchaffe/shakespeare-lines 12.5K 4-line continuation windows + 2.5K per-theme style obama_15k.jsonl 15,000 fivethirtyeight/data BarackObama.csv 6… See the full description on the dataset page: https://huggingface.co/datasets/xinbenlv/gemma-finetune-webgpu.
gemma-finetune-webgpu
Built voice/style instruction-tuned datasets used by the gemma-finetune workshop (May 2026, Immersive Commons). Each row is dolly-15k–shaped:
{"instruction": "...", "context": "...", "response": "...", "category": "..."}Files
Reproducible from the build script in the upstream repo:
git clone https://github.com/RayyanZahid/gemma-finetune
cd gemma-finetune
python data/build_voice_dataset.py --allThe build script uses random.Random(42); output is deterministic.
Use with Gemma fine-tuning
python templates/finetune.py --user me --dataset data/obama_15k.jsonl --out-dir runsLicense
MIT for the derived JSONL. Upstream sources: Project Gutenberg works (Twain) are public domain; fschlatt/trump-tweets is CC0; FiveThirtyEight tweet repo is open; benchaffe/shakespeare-lines is public domain Shakespeare.
