kaushik-harsh-99/Code-Language-Classification
Programming Language Classification Dataset A large-scale, balanced dataset for programming language identification from source code snippets. Overview This dataset contains 1.664 million cleaned and labeled source code samples across 16 programming languages, specifically designed for programming language classification and identification tasks. Unlike many code datasets that are primarily built for code generation or retrieval, this dataset was curated… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Code-Language-Classification.
383
1---2language:3- en4 5license: mit6 7task_categories:8- text-classification9 10 11tags:12- code13- source-code14- programming-languages15- language-identification16- language-classification17- code-classification18- machine-learning19- classification20- benchmark21- code-intelligence22- software-engineering23 24pretty_name: Programming Language Classification Dataset25 26size_categories:27- 1M<n<10M28---29 30# Programming Language Classification Dataset31 32A large-scale, balanced dataset for programming language identification from source code snippets.33 34## Overview35 36This dataset contains **1.664 million cleaned and labeled source code samples** across **16 programming languages**, specifically designed for programming language classification and identification tasks.37 38Unlike many code datasets that are primarily built for code generation or retrieval, this dataset was curated specifically for language classification. Significant effort was invested in deduplication, quality filtering, class balancing, and evaluation split construction to create a reliable benchmark for training and evaluating language identification models.39 40### Key Features41 42- 1.664 million labeled samples43- 16 programming languages44- Class-balanced dataset45- Multi-stage deduplication pipeline46- Quality-filtered samples47- Fixed train/validation/test splits48- Suitable for both traditional ML and neural approaches49 50## Supported Languages51 52| Language |53|-----------|54| Assembly |55| C |56| C++ |57| C# |58| CSS |59| Dart |60| Go |61| HTML |62| Java |63| JavaScript |64| Kotlin |65| Lua |66| Markdown |67| Python |68| Rust |69| TypeScript |70 71## Dataset Statistics72 73| Split | Samples per Language | Total Samples |74|---------|---------:|---------:|75| Train | 100,000 | 1,600,000 |76| Validation | 2,000 | 32,000 |77| Test | 2,000 | 32,000 |78| Total | 104,000 | 1,664,000 |79 80The dataset is fully balanced across all languages to reduce class imbalance and evaluation bias.81 82## Data Format83 84Samples are stored in JSONL format.85 86```json87{88 "content": "def hello_world():\n print('Hello World')",89 "label": "Python"90}91```92 93## Source Dataset94 95This dataset was derived from:96 97https://huggingface.co/datasets/lumees/github-code-2025-language-split98 99The original dataset was language-separated and used as the foundation for further processing, cleaning, filtering, balancing, and dataset construction.100 101## Dataset Construction Pipeline102 103The dataset was created through a multi-stage processing pipeline.104 105### 1. Language-Based Collection106 107Source code was collected from the source dataset and organized by programming language.108 109### 2. Code Chunk Generation110 111Large source files were segmented into smaller code snippets suitable for language classification tasks.112 113This increases sample diversity while making training more efficient.114 115### 3. Multi-Stage Deduplication116 117Several rounds of deduplication were applied to reduce repeated and near-duplicate code fragments.118 119The objective was to improve diversity and reduce memorization of highly repetitive samples.120 121### 4. Quality Filtering122 123A custom filtering pipeline was applied to remove low-quality samples.124 125Filtering stages included:126 127- Low-information sample detection128- Repetitive content removal129- Character distribution analysis130- Symbol-density analysis131- Structural code heuristics132- Corrupted fragment removal133 134Conservative filtering thresholds were intentionally used to preserve valid code across languages with very different syntactic structures such as Assembly, CSS, HTML, Markdown, and Python.135 136### 5. Balanced Sampling137 138Many languages originally contained substantially different numbers of samples.139 140To prevent majority-language bias, a balanced subset was created for every language.141 142Each language contributes exactly:143 144- 100,000 training samples145- 2,000 validation samples146- 2,000 test samples147 148### 6. Split Construction149 150Validation and test sets were sampled independently for each language to ensure balanced and reliable evaluation.151 152## Design Goals153 154The dataset was built with the following objectives:155 156- Balanced language representation157- Reduced duplication158- High sample diversity159- Reliable evaluation splits160- Efficient classifier training161- Support for both classical ML and neural approaches162 163 