datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InstructCoder
Paper |
Code |
Blog
InstructCoder (CodeInstruct): Empowering Language Models to Edit Code
Updates
May 23, 2023: Paper, code and data released.
Overview
InstructCoder is the first dataset designed to adapt LLMs for general code editing. It consists of 114,239 instruction-input-output triplets and covers multiple distinct code editing scenarios, generated by ChatGPT. LLaMA-33B finetuned on InstructCoder performs on par with ChatGPT on a… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/InstructCoder.Evol-Instruct-Code-80k-v1Open Source Implementation of Evol-Instruct-Code as described in the WizardCoder Paper.
Code for the intruction generation can be found on Github as Evol-Teacher.
code_contests_instruct
Dataset Card for "code_contests_instruct"
The deepmind/code_contests dataset formatted as markdown-instruct for text generation training.
There are several different configs. Look at them. Comments:
flesch_reading_ease is computed on the description col via textstat
hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater
min-cols drops all cols except language and text
possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-chat-formatted-generationsself-instruct-starcoder
Self-instruct-starcoder
Summary
Self-instruct-starcoder is a dataset that was generated by prompting starcoder to generate new instructions based on some human-written seed instructions.
The underlying process is explained in the paper self-instruct. This algorithm gave birth to famous machine generated
datasets such as Alpaca and Code Alpaca which are two datasets
obtained by prompting OpenAI text-davinci-003 engine.
Our approach
While our method is… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/self-instruct-starcoder.ds-coder-instruct-v1
Dataset Card for DS Coder Instruct Dataset
DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on
data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python.
The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.self-oss-instruct-sc2-exec-filter-prompt-codes-test-50kdistill_r1_code_evol_instructcode_contest_instruct_cppds-coder-instruct-v2
Dataset Card for DS Coder Instruct v2 Dataset
Changes from v1:
Added WizardLM evol data science samples
Removed R samples from v2
DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on
data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2).
The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.CodeMaster-Phi-Instruct
Code Master Phi is a compiled dataset designed for training Phi3 instruct models. This dataset is focused on code-based data and integrates multiple high-quality sources to ensure a robust training foundation. The sources include:
Replete-AI/code_bagel: A diverse collection of code snippets and examples.
nickrosh/Evol-Instruct-Code-80k-v1: A dataset featuring evolved instructions for code generation tasks.
iamtarun/python_code_instructions_18k_alpaca: A compilation of Python code… See the full description on the dataset page: https://huggingface.co/datasets/thesven/CodeMaster-Phi-Instruct.evol_instruct_code_filtered_39k
Dataset Card for "evol_instruct_code_filtered_38k"
Filtered version of nickrosh/Evol-Instruct-Code-80k-v1, with manual filtering, and automatic filtering based on quality and learning value classifiers.
tiny-codes-instructBase dataset: iamtarun/code_instructions_120k_alpaca
details_Qwen__Qwen2.5-Coder-14B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-14B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-14B-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-14B-Instruct.Vibe-Coding-Instructr2e-code-instruct-qwen3.5
r2e-code-instruct-qwen3.5
Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments.
💡 Browse this dataset in your browser — click the badge above or open
HuggingFaceH4/harbor-visualiser
to inspect every task's spec, instruction, oracle patch, test script, and Dockerfile.
Source repos (39):
adrienverge/yamllint
agronholm/typeguard
alecthomas/voluptuous
andialbrecht/sqlparse
benoitc/gunicorn
bottlepy/bottle
chardet/chardet… See the full description on the dataset page: https://huggingface.co/datasets/Essacheez/r2e-code-instruct-qwen3.5.Code-Instruct-Setsdeepcoder-train-deepcoder-qwen4b-instruct-cont-temp0_6-32k-hsrun_step230-codeonly_truncationVibe-Coding-Instruct-V2
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/CodeDevX/Vibe-Coding-Instruct-V2.gemma4-code-review-instruct
gemma4-code-review-instruct
197K code review examples — 58K with chain-of-thought <think> reasoning traces.
Built to train models that don't just flag issues, but explain their reasoning before delivering a review. Drop-in ready for SFT with any chat model.
Why This Dataset
Most code review datasets give you diff → comment. This one gives you diff → think → comment for 30% of examples — reasoning traces that show how to analyze a diff before writing the review.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/gemma4-code-review-instruct.amenokaku-code-instruct
Amenokaku-Code-Instruct
Update:
2023/12/27データセットに JaxTon , プロになるJava のコードデータ 180 レコードを追加しました。
概要
コードに特化した5.2KのInstructionデータセットです。
データセットに含まれるデータは商用利用できるラインセンスが付与されたプログラミング学習コンテンツから収集、加工し作成しました(英語のコンテンツは日本語に自動翻訳し、翻訳の不自然な箇所を手動で修正)。
また、ライセンスが明記されていない学習コンテンツについては権利者に個別に連絡を取り、本データセットへの掲載の許諾を得ております。
データセット詳細
指示タスクの内訳としてはコード生成(code_generation)が1050レコード、コードの挙動確認(check_code_behavor)が150レコード、コードのバグ修正(code_fix)が4000レコードになります。
詳細な内訳は以下の通りになります。
source name… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/amenokaku-code-instruct.code-review-instruct-critique-revision
Dataset Card for "code-review-instruct-critique-revision"
More Information needed
SynthUI-Code-Instruct-2k-v1Synth UI 🎹
https://www.synthui.design
Dataset details
This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more.
This dataset consists of:
Note: The dataset is seperated into two main parts:
raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-Instruct-2k-v1.HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-dataset-train-generationseval-Qwen3-Coder-30B-A3B-Instruct_16concurrency_openhands_eval_c_terminal-bench-2.0instruct_code_search_net
Dataset Card for "instruct_code_search_net"
More Information needed
details_Qwen__Qwen2.5-Coder-7B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-Coder-7B-Instruct.CodeChat-Instruct-v1
CodeChat-Instruct-v1
CodeChat-Instruct-v1 is a synthetic coding instruction dataset designed for supervised fine-tuning of language models on programming-related conversations. It includes diverse coding tasks such as code review, code improvement, complexity analysis, edge-case discussion, code explanation, library/API usage, refactoring guidance. The dataset is suitable for training coding assistants, educational programming tutors, and general-purpose code LLMs with strong… See the full description on the dataset page: https://huggingface.co/datasets/kd13/CodeChat-Instruct-v1.code-alpaca-instruct-unfilteredThis dataset is HuggingFaceH4/CodeAlpaca_20K unfiltered, removing 36 instances of blatant alignment.
19986 instructions remain.
https://huggingface.co/datasets/HuggingFaceH4/CodeAlpaca_20K/blob/29ba7b7fdf0c55e5435c848cf6bbf9782fef62a6/data/test-00000-of-00001.parquet
https://huggingface.co/datasets/HuggingFaceH4/CodeAlpaca_20K/blob/a123ae447f02484d83c3457438b4422cd8417ad5/data/train-00000-of-00001.parquet
i combined all of these files above into code_alpaca_data.jsonl with parquet2json and ran… See the full description on the dataset page: https://huggingface.co/datasets/ewof/code-alpaca-instruct-unfiltered.Qwen__Qwen2.5-Coder-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-Coder-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-Coder-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-Coder-7B-Instruct-details.
