TGPRO32/Paragon-coding
NOTICE This was done by me, someone with a learning Disability. So please do bare with me when updating this with more working data. Multi-Language Programming Code Dataset A curated dataset of original, non-scraped code examples across 7 programming environments: Python, JavaScript, Node.js, Java, C, C++, and Rust. The dataset ships in two parts that can be used separately or combined: File Rows Description code_dataset.jsonl / .csv 105 Hand-written… See the full description on the dataset page: https://huggingface.co/datasets/TGPRO32/Paragon-coding.
NOTICE
This was done by me, someone with a learning Disability. So please do bare with me when updating this with more working data.
Multi-Language Programming Code Dataset
A curated dataset of original, non-scraped code examples across 7 programming environments: Python, JavaScript, Node.js, Java, C, C++, and Rust.
The dataset ships in two parts that can be used separately or combined:
Note on Node.js: Node.js is a JavaScript runtime, not a separate language. It's included as its own split because it exposes different APIs (filesystem,Buffer,process,http, CommonJS modules) than browser-context JavaScript — which is usually what people actually mean by "Node.js code."
Dataset Structure
Both files share the same schema:
code_dataset (105 rows) — Core Concepts
15 categories × 7 languages, one example each: Hello World · Variables and Data Types · Control Flow · Loops · Functions · Arrays and Lists · Dictionaries and Maps · Classes and OOP · Error Handling · File I/O · String Manipulation · Recursion · Sorting Algorithm · Async and Concurrency · Data Structures
code_dataset_large (2,255 rows) — Deep Coverage on 3 Categories
Generated by varying real parameters — algorithm choice, data type, sample values, operation sequences, and identifier names — not by duplicating templates with find-and-replace. Breakdown:
Distribution is balanced across languages (~320–326 rows each) and skews intermediate (1,727) over beginner (528), reflecting the algorithmic focus of this batch. Exact-duplicate rows were checked and removed (~3% collision rate from small-integer arrays landing on the same random sample).
Files
code_dataset.jsonl/code_dataset_large.jsonl— one JSON object per line (recommended fordatasets.load_dataset("json", ...))code_dataset.csv/code_dataset_large.csv— same data, spreadsheet-friendlygenerate_dataset.py— generator for the 105-row core set (add more languages/categories by addingadd(...)calls)generate_batch.py— generator for the 2,255-row deep-coverage set (add more categories/algorithms by extending the template dicts)
Provenance & License
All code was written from scratch (hand-authored for the core set; programmatically templated with varied real parameters for the large set) — nothing was scraped from GitHub or any other source, so there are no third-party license conflicts. Released under MIT — free to use, modify, and redistribute, including for model training.
Known Limitations
- The 2,255-row set currently covers only 3 categories in depth (sorting, data structures, strings). Categories like "Hello World" or "Variables" don't have enough genuine variation to scale the same way — padding them would mean shallow repetition rather than useful diversity.
- For large-scale pretraining, pair this with an established corpus like The Stack or CodeSearchNet.
- Snippets favor clarity/idiom over production hardening (minimal input validation) — they teach the pattern, not production-ready code.
Example Rows
Core set:
{
"id": 1,
"language": "Python",
"category": "Hello World",
"difficulty": "beginner",
"task_description": "Print 'Hello, World!' to the console.",
"code": "print(\"Hello, World!\")",
"explanation": "Python's print() function writes text to standard output."
}Large set:
{
"id": 11,
"language": "Python",
"category": "Sorting Algorithm",
"difficulty": "intermediate",
"task_description": "Sort a 12-element array of ints in ascending order using bubble sort.",
"code": "def bubble_sort(entries):\n n = len(entries)\n ...",
"explanation": "Bubble sort on int data, variable named 'entries', 12 elements."
}Loading
Hugging Face `datasets`:
from datasets import load_dataset
core = load_dataset("json", data_files="code_dataset.jsonl")
large = load_dataset("json", data_files="code_dataset_large.jsonl")Pandas / Kaggle:
import pandas as pd
core = pd.read_csv("code_dataset.csv")
large = pd.read_csv("code_dataset_large.csv")
combined = pd.concat([core, large], ignore_index=True)Suggested Uses
- Fine-tuning a code-explanation or code-generation model
- Few-shot prompting examples for a coding assistant
- Cross-language idiom comparison (e.g., "how does error handling differ between Python and Rust?")
- Algorithm-variant training data (many sorting/data-structure/string examples with controlled, labeled variation)
- A regression-test seed set for code-generation model evals
Roadmap
The large set can be extended the same way to more categories (recursion, OOP, error handling, file I/O, async) by adding template functions to generate_batch.py — happy to keep scaling this up on request.
