datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodeGen-Diverse-5K
CodeGen-Diverse-5K: Broad Coverage for Competitive Programming
Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset)
Dataset Description
CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions.
Key Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.leetcode-codegen-cpp
LeetCode Code-Gen Dataset — C++
4025 rows. Given a problem statement, its input/output examples, and a
required algorithm/technique, generate a correct C++ solution.
Part of a 4-language collection built from the same source: see the sibling
Python,
Java,
C++, and
JavaScript
datasets.
Verification
Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp.leetcode-codegen-javascript
LeetCode Code-Gen Dataset — JavaScript
631 rows. Given a problem statement, its input/output examples, and a
required algorithm/technique, generate a correct JavaScript solution.
Part of a 4-language collection built from the same source: see the sibling
Python,
Java,
C++, and
JavaScript
datasets.
Verification
Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-javascript.llama3.2-java-codegen-90sft-10meta-claude-v1
LLaMA 3.2 Java Code Generation Dataset (90% SFT, 10% Meta Annotated with Claude)
This dataset contains 100,000 examples for Java method generation based on natural language instructions. It is built from the CodeXGLUE text-to-code dataset and designed to support both pure supervised fine-tuning (SFT) and reflection-based meta-learning approaches using Claude 4 Sonnet as the critique model.
🚀 Trained Models
Two models have been trained on this dataset:
SFT Model:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/llama3.2-java-codegen-90sft-10meta-claude-v1.leetcode-codegen-java
LeetCode Code-Gen Dataset — Java
4068 rows. Given a problem statement, its input/output examples, and a
required algorithm/technique, generate a correct Java solution.
Part of a 4-language collection built from the same source: see the sibling
Python,
Java,
C++, and
JavaScript
datasets.
Verification
Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-java.code-generation-sft-100k
Code Generation SFT (100K)
100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions.
Motivation
Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. This… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.LiveCodeBench-CodeGenerationleetcode-codegen-python
LeetCode Code-Gen Dataset — Python
2522 rows. Given a problem statement, its input/output examples, and a
required algorithm/technique, generate a correct Python solution.
Part of a 4-language collection built from the same source: see the sibling
Python,
Java,
C++, and
JavaScript
datasets.
Verification
Every row was extracted, then executed in a sandboxed subprocess against the problem's own stated examples. Only rows that passed all examples are included -- this… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-python.CodeGen-Deep-5K
CodeGen-Deep-5K: Deep Reasoning for Competitive Programming
Part of the CodeGen suite | CodeGen-Diverse-5K (sister dataset)
Dataset Description
CodeGen-Deep-5K is a deep reasoning dataset designed for training code generation models with enhanced problem-solving capabilities. Unlike traditional datasets, this generates multiple distinct solutions for each problem, providing varied reasoning traces and approaches.
Key Statistics
Total samples: 5,000
Unique… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Deep-5K.cve-backport-codegen-dataset
CVE Backport Code Generation Dataset
Per-hunk code generation dataset for CVE security patch backporting, derived from openSUSE Build Service maintenance patches.
Task
Given a region of vulnerable source code and a description of the upstream CVE fix, the model outputs the fixed version of the code. A programmatic diff then produces the final patch. This plays to LLM strengths in code completion and avoids format-sensitivity issues with direct diff generation.… See the full description on the dataset page: https://huggingface.co/datasets/openSUSE/cve-backport-codegen-dataset.codegen-flutter-v1codegen_kodcode_lc2k_taco_merged
Dataset Card for Dataset Name
Merged likaixin/TACO-verified, Leon-Leee/LeetCodeDataset_rectified, and kodCode/KodCode-Light-RL-10K
Dataset Details
Dataset Description
Curated by: Leon (Me)
Funded by [optional]: AIGCode/Koting Intelligence
Language(s) (NLP): English
License: MIT (following GURU-92K)
Dataset Sources [optional]
Repository: stay tuned
Paper [optional]: stay tuned
Uses
Direct Use
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Leon-Leee/codegen_kodcode_lc2k_taco_merged.code-general-fonction-publique
Code général de la fonction publique, non-instruct (11-12-2023)
This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for legal practice.
Fine-tuning is the process of adapting a pre-trained model to perform specific tasks or cater to particular domains. It involves adjusting the model's parameters through a further round of training on task-specific or domain-specific data. While conventional fine-tuning strategies involve… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-general-fonction-publique.ry_manim_codegen
RY Manim Code Generation
This dataset contains queries and responses for ideal Manim code. The input and output have been preprocessed to keep context clean.
Format:
Jsonl format used instead of Json for better efficiency.
{"query": "<a user's question about a topic>", "output": "<LLM's manim code>"}
Citations:
This repository includes data sourced and processed from the following datasets:… See the full description on the dataset page: https://huggingface.co/datasets/RY7Games/ry_manim_codegen.code_generationCodeGeneration-IQuest
CodeGeneration-IQuest
Execution-based Python code-generation prompts for reinforcement-learning post-training, in the verl rule-reward schema. Each row is a single-turn competitive-programming problem whose reward is computed by executing the model's program against a hidden test suite — a program passes only if every case matches. The collection unifies two execution-scorable sources (Code-Contests-O and DeepCoder) and then difficulty-filters them ("goldilocks", see below) so… See the full description on the dataset page: https://huggingface.co/datasets/OctoReasoner/CodeGeneration-IQuest.prompt-sensitivity-codegen
Prompt Sensitivity in Few-Shot Code Generation Dataset
This dataset contains the full generated-code outputs and pass/fail outcomes used in
our prompt sensitivity study across model families, benchmarks, perturbation axes,
and k-shot settings.
Dataset summary
Rows: 240000
Models: claude-sonnet-4, gemini-2.5-flash, gpt-4o, llama-3.3-70b, qwen2.5-coder-3b
Benchmarks: humaneval, mbpp
Axes: order, phrasing, style
k-shot values: 0, 1, 2, 3
Hugging Face repo:… See the full description on the dataset page: https://huggingface.co/datasets/daksh76/prompt-sensitivity-codegen.prompt-sensitivity-codegen
Anonymous Prompt Sensitivity Dataset
This package contains model generations and evaluation outcomes for an anonymized
submission on prompt sensitivity in few-shot code generation.
What is included
prompt_sensitivity_dataset.jsonl: one row per generated sample
prompt_sensitivity_dataset.csv: tabular view of the same rows
prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available
prompt_variant_spec.json: machine-readable description of the prompt… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.code_generation
Visualization of Code Generation Task Cases Samples
Check dataset samples visualization by viewing Dataset Viewer.
The sampling procedure is guided by the Elo distribution introduced in our method.
Original dataset is release_v5 of livecodebench/code_generation_lite from hugging face.
samples/origin: 879/880
License
This repository is licensed under the Apache License 2.0
