datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Evol-Instruct-Code-80k-v1Open Source Implementation of Evol-Instruct-Code as described in the WizardCoder Paper.
Code for the intruction generation can be found on Github as Evol-Teacher.
code_instructions_122k_alpaca_stylereact-code-instructions
React Code Instructions
Popular Queries
Number of instructions by Model
Unnested Messages
Instructions Added Per Day
Dataset of Claude Artifact esque React Apps generated by Llama 3.1 70B, Llama 3.1 405B, and Deepseek Chat V3.
Examples
Virtual Fitness Trainer Website
LinkedIn Clone
iPhone Calculator
Chipotle Waitlist
Apple Store
python-code-instructions-85k
Python Code Instructions - 85K
Instruction-tuning dataset of Python functions paired with short natural-language instructions derived from repository docstrings.
What changed in this release
This release keeps the original public rows and format, but makes the dataset easier to use responsibly:
exact duplicate rows were removed again using normalized instruction + output hashing
deterministic train, validation, and test splits were added
the dataset card now documents… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/python-code-instructions-85k.Code-Instruct-Setsamenokaku-code-instruct
Amenokaku-Code-Instruct
Update:
2023/12/27データセットに JaxTon , プロになるJava のコードデータ 180 レコードを追加しました。
概要
コードに特化した5.2KのInstructionデータセットです。
データセットに含まれるデータは商用利用できるラインセンスが付与されたプログラミング学習コンテンツから収集、加工し作成しました(英語のコンテンツは日本語に自動翻訳し、翻訳の不自然な箇所を手動で修正)。
また、ライセンスが明記されていない学習コンテンツについては権利者に個別に連絡を取り、本データセットへの掲載の許諾を得ております。
データセット詳細
指示タスクの内訳としてはコード生成(code_generation)が1050レコード、コードの挙動確認(check_code_behavor)が150レコード、コードのバグ修正(code_fix)が4000レコードになります。
詳細な内訳は以下の通りになります。
source name… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/amenokaku-code-instruct.codex-code-instructions-300
Codex Multilingual Code Implementation & Repair 300
A 300-record synthetic coding dataset generated with a Codex-family system and
organized around implementation, repair, edge-case handling, and
already-correct code review tasks.
The exact Codex model/version and original generation configuration could not
be recovered from the available provenance records.
The recovered final dataset is exactly the union of six corrected 50-record
batches.
Dataset Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/codex-code-instructions-300.code-alpaca-instruct-unfilteredThis dataset is HuggingFaceH4/CodeAlpaca_20K unfiltered, removing 36 instances of blatant alignment.
19986 instructions remain.
https://huggingface.co/datasets/HuggingFaceH4/CodeAlpaca_20K/blob/29ba7b7fdf0c55e5435c848cf6bbf9782fef62a6/data/test-00000-of-00001.parquet
https://huggingface.co/datasets/HuggingFaceH4/CodeAlpaca_20K/blob/a123ae447f02484d83c3457438b4422cd8417ad5/data/train-00000-of-00001.parquet
i combined all of these files above into code_alpaca_data.jsonl with parquet2json and ran… See the full description on the dataset page: https://huggingface.co/datasets/ewof/code-alpaca-instruct-unfiltered.turkish-code-instructions
Turkish Code Instructions
Turkish instruction-code pairs for LLM fine-tuning.
Code-Evol-Instruct-OSS
Code-Evol-Instruct-OSS
Summary
Code-Evol-Instruct-OSS is a dataset that was generated with Code Evol-Instruct by prompting open-souce LLMs, WizardLM-13B-v1.2 and WizardCoder-34B-Python.
The underlying process is explained in the paper code-evol-instruct. This algorithm gave birth to famous open-souce code LLMs, WizardCoder-Family.
Our approach
We did not use any closed-source LLMs.
Our seed dataset is sourced from self-instruct-starcoder.
We leverage the… See the full description on the dataset page: https://huggingface.co/datasets/CodeResearch/Code-Evol-Instruct-OSS.kenyan-code-switch-instruct-50k
🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs)
A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules.
Dataset Summary
Total Samples: 50,000 instruction-response pairs
train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.Instruction-Fusion-Code-v1
Instruction Fusion Dataset
Aproximately 100k data samples for code generation.
Fused using seeds from evol-codealpaca-v1.
Generator: gpt-4-1106-preview
Fine-tune
Use 'instruction' and 'output' for SFT.
'prompt1' and prompt2' are the seed instructions used for instruction fusion.
citation
If you use this dataset, please cite our paper
@article{guo2023instruction,
title={Instruction fusion: Advancing prompt evolution through hybridization},
author={Guo… See the full description on the dataset page: https://huggingface.co/datasets/Pasta009/Instruction-Fusion-Code-v1.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-details.qwen3.8-targeted-code-instructions-350
Qwen3.8 Max Targeted Bug Detection & Repair 350
A 350-record synthetic targeted programming dataset generated with Qwen3.8 Max
and subsequently subjected to a full semantic correction audit with
ChatGPT / OpenAI.
The dataset emphasizes correctness judgment, bug detection, debugging, repair,
and tightly constrained programming tasks.
Dataset Summary
The publication artifact contains 350 unique records using the schema:
{
"instruction": "...",
"input": "..."… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-targeted-code-instructions-350.unnatural_code_instructions_20M_tokens_separateunnatural_code_instructions_20MEpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.code-instruction-turkishcode_instruct_alpaca_vicuna_wizardlm_56k_backupBackup of code_instruct_alpaca_vicuna_wizardlm used in rombodawg/MegaCodeTraining112k
Link to the combined dataset bellow
https://huggingface.co/datasets/rombodawg/MegaCodeTraining112k
asl-code-sonnet35-instructSources:
https://huggingface.co/datasets/Augmentation-Scaling-Laws/code_claude3.5_sonnet_10000
https://huggingface.co/datasets/Augmentation-Scaling-Laws/magpie_code_claude3.5_sonnet_10000
https://huggingface.co/datasets/Augmentation-Scaling-Laws/webinstruct_code_claude3.5_sonnet_10000
converted to sharegpt
EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-details.EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.Laravel-13x-Code-InstructionsEpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-details.qwen3.8-code-instructions-350
Qwen3.8 Max Python-Weighted Code Instructions 350
A 350-record synthetic programming instruction dataset generated with
Qwen3.8 Max and reviewed with ChatGPT 5.6 Sol High.
The dataset was designed as a Python-weighted mixed-programming set. The
recovered filename and dataset creator recollection indicate an intended
distribution of approximately 60% Python, although the exact language
distribution was not independently reconstructed from the final artifact.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-code-instructions-350.EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-details
Dataset Card for Evaluation run of EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy
Dataset automatically created during the evaluation run of model EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-details.Instructions_Code_v1
