ISLAM-PO/arabic-to-code-8-langs-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and language-specific quality limits. Viewer note: default is a lightweight preview; select full to load the complete corpus. Current Hub Validation Status Repository claim: 3,000,000 records Dataset Server indexed rows: 1,239,045 Dataset Server estimate: 1,995,159 The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.
Dataset evaluation: See `EVALUATION.md` for schema checks, indexing status, and language-specific quality limits. Viewer note:defaultis a lightweight preview; selectfullto load the complete corpus.
Current Hub Validation Status
- Repository claim: 3,000,000 records
- Dataset Server indexed rows: 1,239,045
- Dataset Server estimate: 1,995,159
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive record count.
Arabic-to-Code Dataset - 8 Languages
Train an Arabic-speaking code model with a raw target of 3,000,000 Arabic instruction-to-code pairs across 8 programming languages and 40 JSONL files (11.55 GB). The current Hub index exposes 1,239,045 rows and estimates 1,995,159 rows; the raw target still requires a full JSONL audit.
1. Contents
- 2. Dataset Summary
- 3. Repository Map
- 4. Languages Table
- 5. Categories Table
- 6. Record Schema
- 7. Loading and Training Usage
- 8. Generation and Reproduction
- 9. Considerations and Limitations
- 10. Contributors
- 11. License
- 12. Citation
- 13. Push to Hugging Face Hub
2. Dataset Summary
3. Repository Map
arabic-to-code-8-langs-3m/
├── README.md
├── LICENSE
├── CITATION.cff
├── .gitattributes
├── 01_Python/ # 375K records
│ ├── 01_basics.jsonl # 75,000
│ ├── 02_functions.jsonl # 75,000
│ ├── 03_oop.jsonl # 75,000
│ ├── 04_algorithms.jsonl # 75,000
│ └── 05_projects.jsonl # 75,000
├── 02_JavaScript/ # 375K - same 5 files
├── 03_Java/ # 375K
├── 04_Cpp/ # 375K
├── 05_HTML_CSS/ # 375K
├── 06_SQL/ # 375K
├── 07_PHP/ # 375K
└── 08_Go/ # 375Kgraph TD
ROOT[arabic-to-code-8-langs-3m/<br/>3M target records - 11.55GB]
ROOT --> PY[01_Python<br/>375K]
ROOT --> JS[02_JavaScript<br/>375K]
ROOT --> JV[03_Java<br/>375K]
ROOT --> CP[04_Cpp<br/>375K]
ROOT --> WEB[05_HTML_CSS<br/>375K]
ROOT --> SQ[06_SQL<br/>375K]
ROOT --> PH[07_PHP<br/>375K]
ROOT --> GO[08_Go<br/>375K]
PY --> F1[basics/functions/oop/algorithms/projects]
F1 --> REC[JSON record<br/>instruction - code - explanation]pie title Records by language (375K each)
"Python" : 375000
"JavaScript" : 375000
"Java" : 375000
"C++" : 375000
"HTML/CSS" : 375000
"SQL" : 375000
"PHP" : 375000
"Go" : 3750004. Languages Table (8 folders)
5. Categories Table (5 files per language)
Per language; multiply by 8 for the global total (each file = 75,000 per language = 600,000 globally).
6. Record Schema
Example (shortened):
{
"id": "python-basics-000001",
"language": "python",
"category": "basics",
"instruction": "اكتب كود بايثون كامل وقابل للتشغيل يقوم بالمهمة التالية: حساب مضروب عدد...",
"code": "# مثال 1 ...\ndef solve_omar(...):\n ... \nif __name__ == \"__main__\":\n main()\n",
"explanation": "هذا الحل بلغة بايثون يشرح ... التعقيد الزمني ... حالات اختبار ...",
"difficulty": "متوسط"
}Python samples were executed during validation (returncode 0).
7. Loading and Training Usage
7.1 Stream the full dataset (recommended)
from datasets import load_dataset
ds = load_dataset("ISLAM-PO/arabic-to-code-8-langs-3m", "full", split="train", streaming=True)
print(next(iter(ds)))
py = ds.filter(lambda x: x["language"] == "python")7.2 One language only
from datasets import load_dataset
ds_py = load_dataset("json", data_files={"train": "hf://datasets/ISLAM-PO/arabic-to-code-8-langs-3m/01_Python/*.jsonl"}, split="train", streaming=True)7.3 Instruction-tuning format (for training your Arabic code model)
def to_prompt(row):
return {
"prompt": f"المستخدم: {row['instruction']}\nالمساعد:\n",
"completion": f"{row['code']}\n\n# الشرح:\n{row['explanation']}",
}
# Use with TRL SFTTrainer or axolotl / unsloth chat template7.4 Fine-tune sketch (Transformers + TRL)
# pip install transformers datasets trl peft
from datasets import load_dataset
from trl import SFTTrainer
ds = load_dataset("json", data_files={"train": "hf://datasets/ISLAM-PO/arabic-to-code-8-langs-3m/01_Python/01_basics.jsonl"}, split="train")
# map with to_prompt(), then SFTTrainer(..., dataset_text_field="prompt")8. Generation and Reproduction
Generated by generate_code_dataset.py (task pools x per-language code builders x Arabic padding).
python generate_code_dataset.py --demo # 4,000 records, ~11 MB
python generate_code_dataset.py --full --per-language 375000 --total-gb 109. Considerations and Limitations
10. Contributors
11. License
CC-BY-4.0. Commercial use allowed with attribution. See LICENSE.
12. Citation
@dataset{arabic_to_code_2026,
title = {Arabic-to-Code Dataset: 8 Languages, 3M target Records},
author = {ISLAM-PO and Contributors},
year = {2026},
publisher = {Hugging Face},
version = {1.0.0},
url = {https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m},
note = {3,000,000 records, 40 JSONL files, 11.55 GB, CC-BY-4.0}
}13. Push to Hugging Face Hub
pip install huggingface_hub datasets
huggingface-cli login
cd arabic-to-code-8-langs-3m
git init; git lfs install; git lfs track "*.jsonl"
huggingface-cli repo create arabic-to-code-8-langs-3m --type dataset --yes
git remote add origin https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m
git add README.md LICENSE CITATION.cff .gitattributes
git commit -m "docs: code dataset card"; git push origin main
git add 01_Python 02_JavaScript; git commit -m "data: batch 1"; git push origin main
# repeat for remaining languagesLast updated: 2026-09-03 | Version: 1.0.0 | Status: complete generation target; verify indexed count before publication / 11.55 GB
