lylybig8/routing_analysis-finetuning-data
routing_analysis finetuning data A mirror of routing_analysis/finetuning/data/: the JSONL training splits used in the multilingual MoE language-expansion experiments, plus the generator scripts, runners, manifests and logs that produced them. Layout Path Contents Size splits/ Document-count tiers {lang}_{tier}.jsonl, {lang}_manifest.json, splits_summary.json ~36 GB token_budget_splits/ Token-budget tiers {lang}_{1Btok,1p5Btok,2Btok}.jsonl and… See the full description on the dataset page: https://huggingface.co/datasets/lylybig8/routing_analysis-finetuning-data.
routing_analysis finetuning data
A mirror of routing_analysis/finetuning/data/: the JSONL training splits used in the multilingual MoE language-expansion experiments, plus the generator scripts, runners, manifests and logs that produced them.
Layout
Every JSONL line in splits/ and token_budget_splits/ is {"id": ..., "text": ..., "language": ...}.
Document-count tiers (splits/)
Documents are taken in canonical source order with len(text) >= 100, and each smaller tier is an exact byte prefix of the next larger one (e.g. 50k ⊂ 100k ⊂ 500k). Each {lang}_manifest.json records the source and the bytes / SHA-256 of every tier.
amh_Ethi and yor_Latn come from data/{lang}/train/000_00000.parquet of HuggingFaceFW/fineweb-2 at revision af9c13333eb981300149d5ca60a8e9d659b276b9 and were built by prepare_fineweb2_amh_yor_splits.py (run_prepare_amh_yor_splits.sh).
Token-budget tiers (token_budget_splits/)
Fixed-seed (42) splits for cat_Latn, dan_Latn, fas_Arab, swe_Latn and tam_Taml at 1B / 1.5B / 2B raw tokenizer tokens, built by prepare_token_budget_splits.py. Per-tier document and token counts are in token_budget_splits/token_budget_splits_summary.md; the plan is in TOKEN_BUDGET_SPLITS.md.
Tamil Indic-Rag-Suite (indicrag_suite_tamil/)
The Tamil subset of ai4bharat/Indic-Rag-Suite (946,544 rows) rendered as Question: ... | Answer: ... | Reasoning: ..., split into 936,993 train and 9,551 validation rows by connected wiki_id components, and packed into 4096-token blocks (*_packed4096/, NumPy arrays with a source map).
Download
# One language's document-count tiers
hf download lylybig8/routing_analysis-finetuning-data --repo-type dataset \
--include "splits/amh_Ethi_*" --local-dir finetuning/data
# Everything
hf download lylybig8/routing_analysis-finetuning-data --repo-type dataset \
--local-dir finetuning/dataAbsolute paths inside manifests and logs refer to the machines the files were generated on.
License and sources
- FineWeb-2 and FineWeb2-HQ: ODC-By 1.0, subject to the Common Crawl terms of use.
- Indic-Rag-Suite (
indicrag_suite_tamil/): MIT.
This repository redistributes subsets of these datasets under their original licenses.
