Team Ai
Datasetpublic

lylybig8/routing_analysis-finetuning-data

routing_analysis finetuning data A mirror of routing_analysis/finetuning/data/: the JSONL training splits used in the multilingual MoE language-expansion experiments, plus the generator scripts, runners, manifests and logs that produced them. Layout Path Contents Size splits/ Document-count tiers {lang}_{tier}.jsonl, {lang}_manifest.json, splits_summary.json ~36 GB token_budget_splits/ Token-budget tiers {lang}_{1Btok,1p5Btok,2Btok}.jsonl and… See the full description on the dataset page: https://huggingface.co/datasets/lylybig8/routing_analysis-finetuning-data.

sourceHugging Faceodc-byupdated 14d agoView on Hugging Face
0likes195downloads
Dataset Card

routing_analysis finetuning data

A mirror of routing_analysis/finetuning/data/: the JSONL training splits used in the multilingual MoE language-expansion experiments, plus the generator scripts, runners, manifests and logs that produced them.

Layout

PathContentsSize
splits/Document-count tiers {lang}_{tier}.jsonl, {lang}_manifest.json, splits_summary.json~36 GB
token_budget_splits/Token-budget tiers {lang}_{1Btok,1p5Btok,2Btok}.jsonl and manifests~66 GB
indicrag_suite_tamil/Tamil QA text from Indic-Rag-Suite, grouped train/validation split, 4096-token packed arrays~3.7 GB
tam_1000k/Generator, verifier and token counts for splits/tam_Taml_1000k.jsonlsmall
*.py, run_*.sh, logs/, *.mdSplit generators, runners, run logs and notessmall

Every JSONL line in splits/ and token_budget_splits/ is {"id": ..., "text": ..., "language": ...}.

Document-count tiers (splits/)

Documents are taken in canonical source order with len(text) >= 100, and each smaller tier is an exact byte prefix of the next larger one (e.g. 50k ⊂ 100k ⊂ 500k). Each {lang}_manifest.json records the source and the bytes / SHA-256 of every tier.

LanguageSourceTiers
amh_EthiFineWeb-250k, 100k
cat_LatnFineWeb-250k, 100k, 500k
dan_LatnFineWeb2-HQ50k, 100k, 500k, 1000k
fas_ArabFineWeb2-HQ50k, 100k, 500k
hau_LatnFineWeb-250k, 100k, 500k
kat_GeorFineWeb-250k, 100k, 500k
swe_LatnFineWeb2-HQ50k, 100k, 500k
swh_LatnFineWeb-250k, 100k, 500k
tam_TamlFineWeb-250k, 100k, 500k, 1000k
yor_LatnFineWeb-250k

amh_Ethi and yor_Latn come from data/{lang}/train/000_00000.parquet of HuggingFaceFW/fineweb-2 at revision af9c13333eb981300149d5ca60a8e9d659b276b9 and were built by prepare_fineweb2_amh_yor_splits.py (run_prepare_amh_yor_splits.sh).

Token-budget tiers (token_budget_splits/)

Fixed-seed (42) splits for cat_Latn, dan_Latn, fas_Arab, swe_Latn and tam_Taml at 1B / 1.5B / 2B raw tokenizer tokens, built by prepare_token_budget_splits.py. Per-tier document and token counts are in token_budget_splits/token_budget_splits_summary.md; the plan is in TOKEN_BUDGET_SPLITS.md.

Tamil Indic-Rag-Suite (indicrag_suite_tamil/)

The Tamil subset of ai4bharat/Indic-Rag-Suite (946,544 rows) rendered as Question: ... | Answer: ... | Reasoning: ..., split into 936,993 train and 9,551 validation rows by connected wiki_id components, and packed into 4096-token blocks (*_packed4096/, NumPy arrays with a source map).

Download

bash
# One language's document-count tiers
hf download lylybig8/routing_analysis-finetuning-data --repo-type dataset \
  --include "splits/amh_Ethi_*" --local-dir finetuning/data

# Everything
hf download lylybig8/routing_analysis-finetuning-data --repo-type dataset \
  --local-dir finetuning/data

Absolute paths inside manifests and logs refer to the machines the files were generated on.

License and sources

This repository redistributes subsets of these datasets under their original licenses.