datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
code_contests_instruct
Dataset Card for "code_contests_instruct"
The deepmind/code_contests dataset formatted as markdown-instruct for text generation training.
There are several different configs. Look at them. Comments:
flesch_reading_ease is computed on the description col via textstat
hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater
min-cols drops all cols except language and text
possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.code_instructions_120k_alpaca
Dataset Card for code_instructions_120k_alpaca
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here.
instructional_code-search-net-python
Dataset Card for "instructional_code-search-net-python"
Dataset Summary
This is an instructional dataset for Python.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.python-code-instructions-japanese
Python Code Instructions - Japanese (18K)
Dataset Description
This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions.
Key Features
18,612 entries covering diverse Python programming tasks
Japanese instructions and prompts for code generation
Original English text preserved for reference
Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.python-code-instructions-85k
Python Code Instructions - 85K
Instruction-tuning dataset of Python functions paired with short natural-language instructions derived from repository docstrings.
What changed in this release
This release keeps the original public rows and format, but makes the dataset easier to use responsibly:
exact duplicate rows were removed again using normalized instruction + output hashing
deterministic train, validation, and test splits were added
the dataset card now documents… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/python-code-instructions-85k.code_contest_instruct_cppgemma4-code-review-instruct
gemma4-code-review-instruct
197K code review examples — 58K with chain-of-thought <think> reasoning traces.
Built to train models that don't just flag issues, but explain their reasoning before delivering a review. Drop-in ready for SFT with any chat model.
Why This Dataset
Most code review datasets give you diff → comment. This one gives you diff → think → comment for 30% of examples — reasoning traces that show how to analyze a diff before writing the review.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/gemma4-code-review-instruct.instructional_code-search-net-javacript
Dataset Card for "instructional_code-search-net-javacript"
Dataset Summary
This is an instructional dataset for JavaScript.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-javacript.instructional_code-search-net-java
Dataset Card for "instructional_code-search-net-java"
Dataset Summary
This is an instructional dataset for Java.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-java.SynthUI-Code-Instruct-2k-v1Synth UI 🎹
https://www.synthui.design
Dataset details
This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more.
This dataset consists of:
Note: The dataset is seperated into two main parts:
raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-Instruct-2k-v1.python-code-instructions-18k-alpaca-tr
Python Code Instructions 18K Alpaca (Turkish)
Turkish translation of Python code instruction dataset for code generation tasks.
Dataset Details
Records: 18,610
Language: Turkish
Format: Alpaca-style instruction/input/output
Columns
Column
Description
text
Formatted training text (instruction + input + code output)
instruction
Turkish instruction
input
Optional input/context
output
Python code solution
Example
{
"text":… See the full description on the dataset page: https://huggingface.co/datasets/mrbesher/python-code-instructions-18k-alpaca-tr.instructional_code-search-net-php
Dataset Card for "instructional_code-search-net-php"
Dataset Summary
This is an instructional dataset for PHP.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-php.kenyan-code-switch-instruct-50k
🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs)
A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules.
Dataset Summary
Total Samples: 50,000 instruction-response pairs
train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.Instruct-Python-Code-Turkish
Dataset Card for Instruct-Python-Code-Turkish
Language: Turkish
Dataset Description
The translation was performed using the Google translation model to ensure high-quality, accurate translation.
Dataset Details
Size: ≈5K
Translation tool: Google Translate
Data format: Instruct, Output
adaption-code-oss-instruct-raw-aug
OSS-Instruct Coding Tasks (Augmented)
Coding problems inspired by open-source snippets, with solutions across several languages.
Rows
8,996
Domain
programming
Format
data.parquet, one row per example
Licence
other
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw-aug.adaption-code-oss-instruct-raw
OSS-Instruct Coding Tasks
Coding problems inspired by open-source snippets, with solutions across several languages.
Rows
3,000
Domain
programming
Format
data.parquet, one row per example
Licence
mit
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.
enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw.adaption-code-oss-instruct-raw-aug-e5bca4
OSS-Instruct Coding Tasks (Augmented)
Coding problems inspired by open-source snippets, with solutions across several languages.
Rows
7,000
Domain
programming
Format
data.parquet, one row per example
Licence
other
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw-aug-e5bca4.ruby-code-instructions-80k
Ruby Code Instructions - 80K
Instruction-tuning dataset of Ruby functions or methods paired with short natural-language instructions derived from repository docstrings or inline comments.
What changed in this release
This release keeps the original public rows and format, but makes the dataset easier to use responsibly:
exact duplicate rows were removed again using normalized instruction + output hashing
deterministic train, validation, and test splits were added
the… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/ruby-code-instructions-80k.code-instruct-mixed
Description
Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence.
Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is.
Usage
from datasets import load_dataset
ds = load_dataset("PotatoHD/code-instruct-mixed")
python-github-code-instruct-filtered-5k
Dataset Card for "python-github-code-instruct-filtered-5k"
This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03.
Feedback and additional columns generated through OpenAI and Cohere responses.
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
Bangla-Code-Instruct
🐯 Bangla-Code-Instruct: A Comprehensive Bangla Code Instruction Dataset
Accepted at LREC 2026
Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri
George Mason University, Fairfax, VA, USA
The first large-scale Bangla code instruction dataset (300K examples) for training Code LLMs in Bangla.
⚠️ Note: The dataset will be released after the LREC 2026 conference. Stay tuned!
Overview
Bangla-Code-Instruct is a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Code-Instruct.gene-code-generation-instruct
code-generation-instruct v2
Gate-passed instruction data for code-generation — published when 50 fresh examples cleared the quality bar
Kind: synthetic
Domain: code-generation
Records: 96
Created: 2026-06-20T19:02:16+00:00
SHA-256: a3f6a919356ea6d71f365ea93c8ea06cb7dc19f22cad57210dedb1f327caed90
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7}
Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-code-generation-instruct.instructional_code-search-net-ruby
Dataset Card for "instructional_code-search-net-ruby"
Dataset Summary
This is an instructional dataset for Ruby.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-ruby.instruct_code_cleaning
SFT code dataset building
Contain a list of tasks useful when building a iniitial dataset source:
reverse_translation
Given a history of conversations, what would the human ask next?
reverse_translation_first_round
Suppose you already have a response, the LLM must predict what question does the human asked
clean_code
Given a code snippet, it determines whether its useful and atomic enough to be use for a response by LLM
gen_code_question
Generates a question given a… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/instruct_code_cleaning.Evol-Instruct-Code-80k-v1-koTranslated nickrosh/Evol-Instruct-Code-80k-v1 using nayohan/llama3-instrucTrans-enko-8b.
This is a raw translation dataset. It needs to be filtered for repetitions generated by the model.
deepseek-ai-deepseek-coder-v2-lite-instruct__code-generation-mbpp-mini__019e3b6f7d82
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on code.generation.mbpp-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
5
N Ok
5
Ok Rate
1
Pass At 1
1
Pass At 1 P05
1
Pass At 1 P50
1
Pass At 1 P95
1
Timeout Rate
0
TTFT P50
55.0369
ms
Total P50 Ms
1212.8
Tokens Out Total
807
Run configuration
Model: deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct @ unknown00
Engine: vllm… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__code-generation-mbpp-mini__019e3b6f7d82.cairo_code_instructions_9k
Cairo Programming Language Dataset
Dataset Summary
This dataset contains a collection of instruction-output pairs focused on the Cairo programming language. It is designed to facilitate the training of language models to understand and generate Cairo code based on given instructions. The dataset includes examples of code snippets, explanations, and tasks related to Cairo, making it suitable for code generation and understanding tasks.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/espejelomar/cairo_code_instructions_9k.deepseek-ai-deepseek-coder-v2-lite-instruct__code-generation-humaneval-mini__019e3b6f4f8a
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on code.generation.humaneval-mini (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
N Samples
5
N Ok
5
Ok Rate
1
Pass At 1
1
Pass At 1 P05
1
Pass At 1 P50
1
Pass At 1 P95
1
Timeout Rate
0
TTFT P50
57.5723
ms
Total P50 Ms
1244.1407
Tokens Out Total
789
Run configuration
Model: deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct @ unknown00
Engine:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__code-generation-humaneval-mini__019e3b6f4f8a.
