Team Ai
Datasetpublic

ISLAM-PO/arabic-to-code-8-langs-3m

Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and language-specific quality limits. Viewer note: default is a lightweight preview; select full to load the complete corpus. Current Hub Validation Status Repository claim: 3,000,000 records Dataset Server indexed rows: 1,239,045 Dataset Server estimate: 1,995,159 The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.

sourceHugging Facecc-by-4.0updated 15d agoView on Hugging Face
0likes267downloads
Dataset Card
Dataset evaluation: See `EVALUATION.md` for schema checks, indexing status, and language-specific quality limits. Viewer note: default is a lightweight preview; select full to load the complete corpus.

Current Hub Validation Status

  • —Repository claim: 3,000,000 records
  • —Dataset Server indexed rows: 1,239,045
  • —Dataset Server estimate: 1,995,159

The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive record count.

Arabic-to-Code Dataset - 8 Languages

Train an Arabic-speaking code model with a raw target of 3,000,000 Arabic instruction-to-code pairs across 8 programming languages and 40 JSONL files (11.55 GB). The current Hub index exposes 1,239,045 rows and estimates 1,995,159 rows; the raw target still requires a full JSONL audit.

1. Contents

2. Dataset Summary

AttributeValue
Dataset namearabic-to-code-8-langs-3m
Programming languages8 (Python, JavaScript, Java, C++, HTML/CSS, SQL, PHP, Go)
Files40 *.jsonl files (8 x 5)
Total records3,000,000 raw target; 1,239,045 currently indexed; 1,995,159 estimated
Total size11.55 GB (11,832.2 MB)
FormatJSONL, UTF-8, one record per line
Content languageArabic instructions + Arabic code comments + complete runnable code
Suggested tasksArabic-to-code generation, instruction tuning, code explanation in Arabic, multilingual code LLM
LicenseCC-BY-4.0
Loadingload_dataset with streaming=True (recommended)

3. Repository Map

text
arabic-to-code-8-langs-3m/
├── README.md
├── LICENSE
├── CITATION.cff
├── .gitattributes
├── 01_Python/          # 375K records
│   ├── 01_basics.jsonl      # 75,000
│   ├── 02_functions.jsonl   # 75,000
│   ├── 03_oop.jsonl         # 75,000
│   ├── 04_algorithms.jsonl  # 75,000
│   └── 05_projects.jsonl    # 75,000
├── 02_JavaScript/      # 375K - same 5 files
├── 03_Java/            # 375K
├── 04_Cpp/             # 375K
├── 05_HTML_CSS/        # 375K
├── 06_SQL/             # 375K
├── 07_PHP/             # 375K
└── 08_Go/              # 375K
mermaid
graph TD
    ROOT[arabic-to-code-8-langs-3m/<br/>3M target records - 11.55GB]
    ROOT --> PY[01_Python<br/>375K]
    ROOT --> JS[02_JavaScript<br/>375K]
    ROOT --> JV[03_Java<br/>375K]
    ROOT --> CP[04_Cpp<br/>375K]
    ROOT --> WEB[05_HTML_CSS<br/>375K]
    ROOT --> SQ[06_SQL<br/>375K]
    ROOT --> PH[07_PHP<br/>375K]
    ROOT --> GO[08_Go<br/>375K]
    PY --> F1[basics/functions/oop/algorithms/projects]
    F1 --> REC[JSON record<br/>instruction - code - explanation]
mermaid
pie title Records by language (375K each)
    "Python" : 375000
    "JavaScript" : 375000
    "Java" : 375000
    "C++" : 375000
    "HTML/CSS" : 375000
    "SQL" : 375000
    "PHP" : 375000
    "Go" : 375000

4. Languages Table (8 folders)

#FolderLanguageRecordsSizeWhat is inside
101_PythonPython375,000~1,536 MBScripts, functions, OOP classes, algorithms, full apps - all runnable
202_JavaScriptJavaScript375,000~1,403 MBNode.js + browser code, functions, classes, DOM examples
303_JavaJava375,000~1,527 MBFull Main.java programs, collections, OOP
404_CppC++375,000~1,533 MBFull g++-ready programs with STL
505_HTML_CSSHTML/CSS375,000~1,437 MBComplete responsive RTL Arabic pages with JS
606_SQLSQL375,000~1,520 MBCREATE + INSERT + SELECT scripts (Postgres/MySQL)
707_PHPPHP375,000~1,432 MBCLI-ready .php scripts
808_GoGo375,000~1,444 MBgo run-ready programs
Total8 folders—3,000,00011,832.2 MB (11.55 GB)—

5. Categories Table (5 files per language)

Per language; multiply by 8 for the global total (each file = 75,000 per language = 600,000 globally).

#FileRecords/langGlobal (x8)Content (10 task families each)
101_basics.jsonl75,000600,000Sum, factorial, prime check, multiplication table, average, temperature, vowels, reverse string, circle area, leap year
202_functions.jsonl75,000600,000Add, max, filter evens, frequency, merge, email validation, password strength, slugify, age calc, pagination
303_oop.jsonl75,000600,000Student, bank account, car, employee, library, animal inheritance, cart, hospital, course, inventory
404_algorithms.jsonl75,000600,000Bubble sort, binary search, Fibonacci, BFS, Dijkstra, Quicksort, Mergesort, knapsack, KMP, LCS
505_projects.jsonl75,000600,000Task manager, calculator, login system, mini shop, clinic booking, library system, sales dashboard, Arabic chatbot, payroll, portfolio site

6. Record Schema

FieldTypeExampleDescription
idstringpython-basics-000001Unique ID: language-category-number
languagestringpythonOne of 8 languages
categorystringbasicsOne of basics/functions/oop/algorithms/projects
instructionstring (Arabic)اكتب كود بايثون كامل...Arabic task description with N and variable
codestringdef solve_omar...Complete runnable code with Arabic comments
explanationstring (Arabic)هذا الحل بلغة بايثون...Long Arabic explanation + complexity + tests
difficultystringمبتدئ/متوسط/متقدمBeginner / intermediate / advanced (cycled)

Example (shortened):

json
{
  "id": "python-basics-000001",
  "language": "python",
  "category": "basics",
  "instruction": "اكتب كود بايثون كامل وقابل للتشغيل يقوم بالمهمة التالية: حساب مضروب عدد...",
  "code": "# مثال 1 ...\ndef solve_omar(...):\n ... \nif __name__ == \"__main__\":\n    main()\n",
  "explanation": "هذا الحل بلغة بايثون يشرح ... التعقيد الزمني ... حالات اختبار ...",
  "difficulty": "متوسط"
}

Python samples were executed during validation (returncode 0).

7. Loading and Training Usage

7.1 Stream the full dataset (recommended)

python
from datasets import load_dataset
ds = load_dataset("ISLAM-PO/arabic-to-code-8-langs-3m", "full", split="train", streaming=True)
print(next(iter(ds)))
py = ds.filter(lambda x: x["language"] == "python")

7.2 One language only

python
from datasets import load_dataset
ds_py = load_dataset("json", data_files={"train": "hf://datasets/ISLAM-PO/arabic-to-code-8-langs-3m/01_Python/*.jsonl"}, split="train", streaming=True)

7.3 Instruction-tuning format (for training your Arabic code model)

python
def to_prompt(row):
    return {
        "prompt": f"المستخدم: {row['instruction']}\nالمساعد:\n",
        "completion": f"{row['code']}\n\n# الشرح:\n{row['explanation']}",
    }
# Use with TRL SFTTrainer or axolotl / unsloth chat template

7.4 Fine-tune sketch (Transformers + TRL)

python
# pip install transformers datasets trl peft
from datasets import load_dataset
from trl import SFTTrainer
ds = load_dataset("json", data_files={"train": "hf://datasets/ISLAM-PO/arabic-to-code-8-langs-3m/01_Python/01_basics.jsonl"}, split="train")
# map with to_prompt(), then SFTTrainer(..., dataset_text_field="prompt")

8. Generation and Reproduction

Generated by generate_code_dataset.py (task pools x per-language code builders x Arabic padding).

bash
python generate_code_dataset.py --demo                              # 4,000 records, ~11 MB
python generate_code_dataset.py --full --per-language 375000 --total-gb 10
ParameterDefaultMeaning
--per-language375000Records per language (x8 = 3M target total)
--total-gb10Target size; per-record bytes auto-computed

9. Considerations and Limitations

TopicDetails
Synthetic dataTemplate-generated. Great for bootstrapping an Arabic code model, but mix with real repos (e.g. The Stack, CodeAlpaca-Arabic) before production.
CorrectnessPython samples smoke-tested. Other languages follow the same logic but run your own validator for Java/C++/Go/SQL before release.
RepetitionTasks cycle over 10 families per category with varied N/vars/UIDs. Deduplicate + add real-world seeds for final eval.
SecurityNo secrets included. Generated table names contain numeric suffixes only. Never train on private keys.
Size11.55 GB - use streaming or per-language loads.

10. Contributors

RoleNameContribution
Owner & conceptISLAM-POArabic-code training idea, 8 languages + 10 GB requirement
Generation & docsMuse Spark (AI Assistant)generate_code_dataset.py, templates, validation, docs
Code review (open)Open callNative devs: verify Java/C++/Go/SQL samples, add idiomatic patterns

11. License

CC-BY-4.0. Commercial use allowed with attribution. See LICENSE.

12. Citation

bibtex
@dataset{arabic_to_code_2026,
  title   = {Arabic-to-Code Dataset: 8 Languages, 3M target Records},
  author  = {ISLAM-PO and Contributors},
  year    = {2026},
  publisher = {Hugging Face},
  version = {1.0.0},
  url     = {https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m},
  note    = {3,000,000 records, 40 JSONL files, 11.55 GB, CC-BY-4.0}
}

13. Push to Hugging Face Hub

bash
pip install huggingface_hub datasets
huggingface-cli login
cd arabic-to-code-8-langs-3m
git init; git lfs install; git lfs track "*.jsonl"
huggingface-cli repo create arabic-to-code-8-langs-3m --type dataset --yes
git remote add origin https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m
git add README.md LICENSE CITATION.cff .gitattributes
git commit -m "docs: code dataset card"; git push origin main
git add 01_Python 02_JavaScript; git commit -m "data: batch 1"; git push origin main
# repeat for remaining languages

Last updated: 2026-09-03 | Version: 1.0.0 | Status: complete generation target; verify indexed count before publication / 11.55 GB