Team Ai
Datasetpublic

TGPRO32/Paragon-coding

NOTICE This was done by me, someone with a learning Disability. So please do bare with me when updating this with more working data. Multi-Language Programming Code Dataset A curated dataset of original, non-scraped code examples across 7 programming environments: Python, JavaScript, Node.js, Java, C, C++, and Rust. The dataset ships in two parts that can be used separately or combined: File Rows Description code_dataset.jsonl / .csv 105 Hand-written… See the full description on the dataset page: https://huggingface.co/datasets/TGPRO32/Paragon-coding.

sourceHugging Facemitupdated 25d agoView on Hugging Face
2likes183downloads
README.md172 linesDownload Raw Back to root
1---2license: mit3task_categories:4  - text-generation5  - text2text-generation6language:7  - code8tags:9  - code10  - programming-languages11  - python12  - javascript13  - nodejs14  - java15  - c16  - cpp17  - rust18  - sorting-algorithms19  - data-structures20  - synthetic21pretty_name: Multi-Language Programming Code Dataset22size_categories:23  - 1K<n<10K24---25 26# NOTICE27 28This was done by me, someone with a learning Disability. So please do bare with me when updating this with more working data.29 30# Multi-Language Programming Code Dataset31 32A curated dataset of **original, non-scraped** code examples across 7 programming33environments: **Python, JavaScript, Node.js, Java, C, C++, and Rust**.34 35The dataset ships in two parts that can be used separately or combined:36 37| File | Rows | Description |38|---|---|---|39| `code_dataset.jsonl` / `.csv` | 105 | Hand-written "core concepts" set — one clean example per language per concept (Hello World, OOP, error handling, recursion, async, etc.) |40| `code_dataset_large.jsonl` / `.csv` | 2,255 | Template-generated, parameter-varied set covering **sorting algorithms, data structures, and string manipulation** in depth |41 42> **Note on Node.js:** Node.js is a JavaScript *runtime*, not a separate language.43> It's included as its own split because it exposes different APIs (filesystem,44> `Buffer`, `process`, `http`, CommonJS modules) than browser-context JavaScript —45> which is usually what people actually mean by "Node.js code."46 47## Dataset Structure48 49Both files share the same schema:50 51| Column | Description |52|---|---|53| `id` | Unique row identifier (per file) |54| `language` | One of: Python, JavaScript, Node.js, Java, C, C++, Rust |55| `category` | Concept/topic covered |56| `difficulty` | `beginner`, `intermediate`, or `advanced` |57| `task_description` | Natural-language description of the coding task |58| `code` | The code snippet solving the task |59| `explanation` | A short note on the key language feature/idiom used |60 61## `code_dataset` (105 rows) — Core Concepts62 6315 categories × 7 languages, one example each:64Hello World · Variables and Data Types · Control Flow · Loops · Functions ·65Arrays and Lists · Dictionaries and Maps · Classes and OOP · Error Handling ·66File I/O · String Manipulation · Recursion · Sorting Algorithm ·67Async and Concurrency · Data Structures68 69## `code_dataset_large` (2,255 rows) — Deep Coverage on 3 Categories70 71Generated by varying real parameters — algorithm choice, data type, sample72values, operation sequences, and identifier names — **not** by duplicating73templates with find-and-replace. Breakdown:74 75| Category | Rows | What varies |76|---|---|---|77| Sorting Algorithm | 756 | Algorithm (bubble/selection/insertion), data type (int/float), array size (5–20 elements), 3 random samples per config |78| Data Structures | 448 | Stack vs. Queue, data type (int/float), 4 distinct push/pop operation sequences, random values |79| String Manipulation | 1,051 | Operation (palindrome check, reverse, word count, vowel count), 20 distinct test strings, varied function names |80 81Distribution is balanced across languages (~320–326 rows each) and skews82`intermediate` (1,727) over `beginner` (528), reflecting the algorithmic focus83of this batch. Exact-duplicate rows were checked and removed (~3% collision84rate from small-integer arrays landing on the same random sample).85 86## Files87 88- `code_dataset.jsonl` / `code_dataset_large.jsonl` — one JSON object per line (recommended for `datasets.load_dataset("json", ...)`)89- `code_dataset.csv` / `code_dataset_large.csv` — same data, spreadsheet-friendly90- `generate_dataset.py` — generator for the 105-row core set (add more languages/categories by adding `add(...)` calls)91- `generate_batch.py` — generator for the 2,255-row deep-coverage set (add more categories/algorithms by extending the template dicts)92 93## Provenance & License94 95All code was **written from scratch** (hand-authored for the core set;96programmatically templated with varied real parameters for the large set) —97nothing was scraped from GitHub or any other source, so there are no98third-party license conflicts. Released under **MIT** — free to use, modify,99and redistribute, including for model training.100 101## Known Limitations102 103- The 2,255-row set currently covers only 3 categories in depth (sorting,104  data structures, strings). Categories like "Hello World" or "Variables"105  don't have enough genuine variation to scale the same way — padding them106  would mean shallow repetition rather than useful diversity.107- For large-scale pretraining, pair this with an established corpus like108  [The Stack](https://huggingface.co/datasets/bigcode/the-stack) or109  [CodeSearchNet](https://huggingface.co/datasets/code_search_net).110- Snippets favor clarity/idiom over production hardening (minimal input111  validation) — they teach the *pattern*, not production-ready code.112 113## Example Rows114 115**Core set:**116```json117{118  "id": 1,119  "language": "Python",120  "category": "Hello World",121  "difficulty": "beginner",122  "task_description": "Print 'Hello, World!' to the console.",123  "code": "print(\"Hello, World!\")",124  "explanation": "Python's print() function writes text to standard output."125}126```127 128**Large set:**129```json130{131  "id": 11,132  "language": "Python",133  "category": "Sorting Algorithm",134  "difficulty": "intermediate",135  "task_description": "Sort a 12-element array of ints in ascending order using bubble sort.",136  "code": "def bubble_sort(entries):\n    n = len(entries)\n    ...",137  "explanation": "Bubble sort on int data, variable named 'entries', 12 elements."138}139```140 141## Loading142 143**Hugging Face `datasets`:**144```python145from datasets import load_dataset146core = load_dataset("json", data_files="code_dataset.jsonl")147large = load_dataset("json", data_files="code_dataset_large.jsonl")148```149 150**Pandas / Kaggle:**151```python152import pandas as pd153core = pd.read_csv("code_dataset.csv")154large = pd.read_csv("code_dataset_large.csv")155combined = pd.concat([core, large], ignore_index=True)156```157 158## Suggested Uses159 160- Fine-tuning a code-explanation or code-generation model161- Few-shot prompting examples for a coding assistant162- Cross-language idiom comparison (e.g., "how does error handling differ163  between Python and Rust?")164- Algorithm-variant training data (many sorting/data-structure/string165  examples with controlled, labeled variation)166- A regression-test seed set for code-generation model evals167 168## Roadmap169 170The large set can be extended the same way to more categories (recursion,171OOP, error handling, file I/O, async) by adding template functions to172`generate_batch.py` — happy to keep scaling this up on request.