datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiLang-Code-Parser-Dataset
MultiLang Code Parser Dataset (MLCPD)
MultiLang-Code-Parser-Dataset (MLCPD) provides a large-scale, unified dataset of parsed source code across 10 major programming languages, represented under a universal schema that captures syntax, semantics, and structure in a consistent format.
Each entry corresponds to one parsed source file and includes:
Language metadata
Code-level statistics (lines, errors, AST nodes)
Universal Schema JSON (normalized structural representation)
MLCPD… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/MultiLang-Code-Parser-Dataset.code_contests_dp_datasetStackPulse_778K_QnA_Code_dataset
💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset
Dataset Summary
A large-scale Stack Overflow question dataset containing 778,929 unique
questions sampled across 7 years (2015–2022). Each question includes the
raw HTML body, plain-text version, tags, score, view count, answer count, and
a rich set of derived features for immediate ML use.
Collected across 8 sampling runs on Feb 27 2026, deduplicated to
778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.team1_dataset_play_diatonic_codeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/peiran20030408/team1_dataset_play_diatonic_code.reverse_engineering_code_dataset_O2_x64_O2reverse_engineering_code_dataset_O0_arm_O0reverse_engineering_code_dataset_O3_x64_O3reverse_engineering_code_dataset_O2_arm_O2reverse_engineering_code_dataset_O3_arm_O3code-trainer-offsec-datasetreverse_engineering_code_dataset_O2_mips_O2reverse_engineering_code_dataset_O0_x64_O0reverse_engineering_code_dataset_O1_mips_O1AI-Code-Optimization-for-Sustainability-Dataset
AI Code Optimization for Sustainability: Dataset
Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury
📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author
This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency.
The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.reverse_engineering_code_dataset_O1_arm_O1reverse_engineering_code_dataset_O1_x64_O1Python-React-Code-DatasetPython-Security-Code-Datasetreverse_engineering_code_dataset_O0_mips_O0recreated_reverse_engineering_code_dataset_O1_x86_O1pde-llm-eval-code-perturbation-dataset
pde-llm-eval-code-perturbation-dataset
Code Perturbation Dataset (final). 256 PDE solver implementations: 64 base programs (32 physically valid, 32 with an injected bug that still runs to completion) expanded by 4 lexical perturbation conditions each -- original comments, comments removed, comments swapped in from a different implementation, and descriptive identifiers replaced by meaningless placeholders. The perturbations change the lexical surface only; executable behaviour… See the full description on the dataset page: https://huggingface.co/datasets/bermaneh/pde-llm-eval-code-perturbation-dataset.Dr-Zeon-Github-Python-Code-Dataset
Luck Spark 1B - High Quality Code Dataset
The first quality-scored, star-agnostic code dataset for training 1B MoE code models.
Unlike The Stack / CodeParrot that filter by stars, this dataset scores every file by its content (0-10). A 2-star well-documented library scores higher than a 10k-star minified file. Continuously updated by an autonomous bot.
Repo: ahmetggg/luck-spark-1b-code-dataset | Bot: github_to_hf_bot.py | License: Permissive only (MIT / Apache-2.0 / BSD /… See the full description on the dataset page: https://huggingface.co/datasets/ahmetggg/Dr-Zeon-Github-Python-Code-Dataset.titanic_datasetreverse_engineering_code_dataset_O1_x86_O1recreated_reverse_engineering_code_dataset_O3_x86_O3reverse_engineering_code_dataset_O3_x86_O3Code_Vulnerability_Dataset
🔐 Code Vulnerability Dataset (CWE-Enriched)
📌 Overview
This dataset is built from the bstee615/diversevul dataset and enhanced with structured vulnerability intelligence from the MITRE Common Weakness Enumeration (CWE) database.
It provides a rich, machine-readable representation of software vulnerabilities, mapping raw vulnerable code samples to standardized CWE classifications.
The dataset is designed for research and development in:
Vulnerability detection models… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset.dataset_qwen2.5_code_1.5b_grpo_iter0_full_data_miao_0212_2_global_step_70filteredreverse_engineering_code_dataset_O3_mips_O3
