Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01likaixin /InstructCoder Paper | Code | Blog InstructCoder (CodeInstruct): Empowering Language Models to Edit Code Updates May 23, 2023: Paper, code and data released. Overview InstructCoder is the first dataset designed to adapt LLMs for general code editing. It consists of 114,239 instruction-input-output triplets and covers multiple distinct code editing scenarios, generated by ChatGPT. LLaMA-33B finetuned on InstructCoder performs on par with ChatGPT on a… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/InstructCoder.texttext-generation100K<n<1M17 likes9.5k downloads2y agoHugging Face02BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.4k downloads9mo agoHugging Face03ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes366 downloads3y agoHugging Face04wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes254 downloads2y agoHugging Face05ed001 /ds-coder-instruct-v2 Dataset Card for DS Coder Instruct v2 Dataset Changes from v1: Added WizardLM evol data science samples Removed R samples from v2 DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2). The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.tabulartext-generation10K<n<100K13 likes226 downloads3y agoHugging Face06thesven /CodeMaster-Phi-Instruct Code Master Phi is a compiled dataset designed for training Phi3 instruct models. This dataset is focused on code-based data and integrates multiple high-quality sources to ensure a robust training foundation. The sources include: Replete-AI/code_bagel: A diverse collection of code snippets and examples. nickrosh/Evol-Instruct-Code-80k-v1: A dataset featuring evolved instructions for code generation tasks. iamtarun/python_code_instructions_18k_alpaca: A compilation of Python code… See the full description on the dataset page: https://huggingface.co/datasets/thesven/CodeMaster-Phi-Instruct.texttext-generation1M<n<10M0 likes204 downloads2y agoHugging Face07CodeDevX /Vibe-Coding-Instructtexttext-generation1M<n<10M189 likes149 downloads4mo agoHugging Face08liodon-ai /gemma4-code-review-instruct gemma4-code-review-instruct 197K code review examples — 58K with chain-of-thought <think> reasoning traces. Built to train models that don't just flag issues, but explain their reasoning before delivering a review. Drop-in ready for SFT with any chat model. Why This Dataset Most code review datasets give you diff → comment. This one gives you diff → think → comment for 30% of examples — reasoning traces that show how to analyze a diff before writing the review.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/gemma4-code-review-instruct.texttext-generation100K<n<1M4 likes124 downloads4mo agoHugging Face09JulianAT /SynthUI-Code-Instruct-2k-v1Synth UI 🎹 https://www.synthui.design Dataset details This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more. This dataset consists of: Note: The dataset is seperated into two main parts: raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-Instruct-2k-v1.texttext-generation1K<n<10K0 likes94 downloads2y agoHugging Face10kd13 /CodeChat-Instruct-v1 CodeChat-Instruct-v1 CodeChat-Instruct-v1 is a synthetic coding instruction dataset designed for supervised fine-tuning of language models on programming-related conversations. It includes diverse coding tasks such as code review, code improvement, complexity analysis, edge-case discussion, code explanation, library/API usage, refactoring guidance. The dataset is suitable for training coding assistants, educational programming tutors, and general-purpose code LLMs with strong… See the full description on the dataset page: https://huggingface.co/datasets/kd13/CodeChat-Instruct-v1.texttext-generation10K<n<100K1 likes77 downloads4mo agoHugging Face11Maxyelow /kenyan-code-switch-instruct-50k 🇰🇪 Kenyan Code-Switching & Sheng Multi-Task Instruction Dataset (50,000 Pairs) A standardized, multi-task instruction-tuning dataset engineered to teach Large Language Models (e.g. Llama 3, Mistral, Gemma, Qwen) to understand and generate authentic Kenyan Code-Switching (Sheng, Technical Swahili-English Blend) with rigorous adherence to Bantu morphotactic rules. Dataset Summary Total Samples: 50,000 instruction-response pairs train.jsonl: 45,000 pairs (90%)… See the full description on the dataset page: https://huggingface.co/datasets/Maxyelow/kenyan-code-switch-instruct-50k.texttext-generation10K<n<100K0 likes70 downloads10d agoHugging Face12erythropygia /Instruct-Python-Code-Turkish Dataset Card for Instruct-Python-Code-Turkish Language: Turkish Dataset Description The translation was performed using the Google translation model to ensure high-quality, accurate translation. Dataset Details Size: ≈5K Translation tool: Google Translate Data format: Instruct, Output texttext-generation1K<n<10K1 likes64 downloads2y agoHugging Face13rodriguescarson /adaption-code-oss-instruct-raw-aug OSS-Instruct Coding Tasks (Augmented) Coding problems inspired by open-source snippets, with solutions across several languages. Rows 8,996 Domain programming Format data.parquet, one row per example Licence other Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw-aug.texttext-generation1K<n<10K0 likes60 downloads12d agoHugging Face14rodriguescarson /adaption-code-oss-instruct-raw OSS-Instruct Coding Tasks Coding problems inspired by open-source snippets, with solutions across several languages. Rows 3,000 Domain programming Format data.parquet, one row per example Licence mit Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded. enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw.tabulartext-generation1K<n<10K0 likes58 downloads12d agoHugging Face15rodriguescarson /adaption-code-oss-instruct-raw-aug-e5bca4 OSS-Instruct Coding Tasks (Augmented) Coding problems inspired by open-source snippets, with solutions across several languages. Rows 7,000 Domain programming Format data.parquet, one row per example Licence other Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw-aug-e5bca4.texttext-generation1K<n<10K0 likes58 downloads12d agoHugging Face16PotatoHD /code-instruct-mixed Description Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence. Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is. Usage from datasets import load_dataset ds = load_dataset("PotatoHD/code-instruct-mixed") texttext-generation100K<n<1M0 likes51 downloads4mo agoHugging Face17jtatman /python-github-code-instruct-filtered-5k Dataset Card for "python-github-code-instruct-filtered-5k" This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03. Feedback and additional columns generated through OpenAI and Cohere responses. texttext-generation1K<n<10K7 likes50 downloads2y agoHugging Face18md-nishat-008 /Bangla-Code-Instruct 🐯 Bangla-Code-Instruct: A Comprehensive Bangla Code Instruction Dataset Accepted at LREC 2026 Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri George Mason University, Fairfax, VA, USA The first large-scale Bangla code instruction dataset (300K examples) for training Code LLMs in Bangla. ⚠️ Note: The dataset will be released after the LREC 2026 conference. Stay tuned! Overview Bangla-Code-Instruct is a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Code-Instruct.texttext-generation100K<n<1M0 likes46 downloads6mo agoHugging Face19Gene829 /gene-code-generation-instruct code-generation-instruct v2 Gate-passed instruction data for code-generation — published when 50 fresh examples cleared the quality bar Kind: synthetic Domain: code-generation Records: 96 Created: 2026-06-20T19:02:16+00:00 SHA-256: a3f6a919356ea6d71f365ea93c8ea06cb7dc19f22cad57210dedb1f327caed90 Pipeline: v2.0.0 Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7} Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-code-generation-instruct.text-generationn<1K0 likes46 downloads4mo agoHugging Face20Yobitel /deepseek-ai-deepseek-coder-v2-lite-instruct__llm-quality-persona-consistency-mini__019e3b6fdda4 deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on llm.quality.persona-consistency-mini (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 5 N Ok 5 Ok Rate 1 Persona Consistency Mean 0.88 Accuracy 0.88 Persona Consistency P50 0.8 Persona Consistency P95 1 Accuracy P05 0.8 Accuracy P50 0.8 Accuracy P95 1 Drift Rate 0.6 Mean Drift Turn 2.6667 TTFT P50 57.5324 ms Total P50 Ms 2672.8064… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__llm-quality-persona-consistency-mini__019e3b6fdda4.text-generationn<1K0 likes40 downloads5mo agoHugging Face21syntaxsynth /instruct_code_cleaning SFT code dataset building Contain a list of tasks useful when building a iniitial dataset source: reverse_translation Given a history of conversations, what would the human ask next? reverse_translation_first_round Suppose you already have a response, the LLM must predict what question does the human asked clean_code Given a code snippet, it determines whether its useful and atomic enough to be use for a response by LLM gen_code_question Generates a question given a… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/instruct_code_cleaning.texttext-generation10K<n<100K1 likes35 downloads3y agoHugging Face22nayohan /Evol-Instruct-Code-80k-v1-koTranslated nickrosh/Evol-Instruct-Code-80k-v1 using nayohan/llama3-instrucTrans-enko-8b. This is a raw translation dataset. It needs to be filtered for repetitions generated by the model. texttext-generation10K<n<100K1 likes33 downloads2y agoHugging Face23Yobitel /deepseek-ai-deepseek-coder-v2-lite-instruct__llm-inference-chatbot-short__019e3b6f22ca deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit TTFT P50 74.4268 ms TTFT P99 521.1922 ms TPOT P50 21.0371 ms TPOT P99 23.458 ms Total P50 Ms 2661.4077 Total P99 Ms 3136.5972 Req Per S Passing 1.0172 Req Per S All 1.1559 Compliance Rate 0.88 Ok Rate 1 Throughput Tok Per S 134.316 Power Avg W 808.3085 Power Peak W 854.62… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__llm-inference-chatbot-short__019e3b6f22ca.text-generationn<1K0 likes33 downloads5mo agoHugging Face24Yobitel /deepseek-ai-deepseek-coder-v2-lite-instruct__code-generation-mbpp-mini__019e3b6f7d82 deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on code.generation.mbpp-mini (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 5 N Ok 5 Ok Rate 1 Pass At 1 1 Pass At 1 P05 1 Pass At 1 P50 1 Pass At 1 P95 1 Timeout Rate 0 TTFT P50 55.0369 ms Total P50 Ms 1212.8 Tokens Out Total 807 Run configuration Model: deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct @ unknown00 Engine: vllm… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__code-generation-mbpp-mini__019e3b6f7d82.text-generationn<1K0 likes32 downloads5mo agoHugging Face25Yobitel /deepseek-ai-deepseek-coder-v2-lite-instruct__llm-quality-factual-mini__019e3b6f9bdd deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on llm.quality.factual-mini (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 10 N Ok 10 Ok Rate 1 Accuracy 1 Accuracy P05 1 Accuracy P50 1 Accuracy P95 1 TTFT P50 45.0118 ms Total P50 Ms 359.225 Tokens Out Total 373 Run configuration Model: deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct @ unknown00 Engine: vllm vunknownQuantization:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__llm-quality-factual-mini__019e3b6f9bdd.text-generationn<1K0 likes30 downloads5mo agoHugging Face26Yobitel /deepseek-ai-deepseek-coder-v2-lite-instruct__code-generation-humaneval-mini__019e3b6f4f8a deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on code.generation.humaneval-mini (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 5 N Ok 5 Ok Rate 1 Pass At 1 1 Pass At 1 P05 1 Pass At 1 P50 1 Pass At 1 P95 1 Timeout Rate 0 TTFT P50 57.5723 ms Total P50 Ms 1244.1407 Tokens Out Total 789 Run configuration Model: deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct @ unknown00 Engine:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__code-generation-humaneval-mini__019e3b6f4f8a.text-generationn<1K0 likes29 downloads5mo agoHugging Face27codezakh /EFAGen-Llama-3.1-8B-Instruct-Training-DataPaper Link The training data used for the final version of EFAGen-Llama-3.1-8B-Instruct. The data is in Alpaca format and can be used with Llama-Factory (check dataset_info.json). texttext-generation1K<n<10K1 likes28 downloads1y agoHugging Face28Yobitel /microsoft-phi-3-5-mini-instruct__code-generation-humaneval-mini__019e3b5d8c8b microsoft/Phi-3.5-mini-instruct on code.generation.humaneval-mini (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 5 N Ok 5 Ok Rate 1 Pass At 1 1 Pass At 1 P05 1 Pass At 1 P50 1 Pass At 1 P95 1 Timeout Rate 0 TTFT P50 12.396 ms Total P50 Ms 974.8963 Tokens Out Total 1166 Run configuration Model: microsoft/Phi-3.5-mini-instruct @ unknown00 Engine: vllm vunknown Quantization:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/microsoft-phi-3-5-mini-instruct__code-generation-humaneval-mini__019e3b5d8c8b.text-generationn<1K0 likes28 downloads5mo agoHugging Face29HachiML /amenokaku-code-instruct-python-mit-450kunishou/amenokaku-code-instructを以下の条件で絞り込んだものです。 MITライセンス (licence: 'MIT') python (source: ['gasyori_100_knocks', 'datascience_100_knocks_python', 'bifi', 'python_for_begginers_solve_50_exercises', 'nlp_100_knocks') source: 'bifi'をランダムに100件に絞り込み texttext-generationn<1K0 likes25 downloads2y agoHugging Face30Yobitel /meta-llama-llama-3-1-70b-instruct__code-generation-humaneval-mini__019e3b935675 meta-llama/Llama-3.1-70B-Instruct on code.generation.humaneval-mini (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 5 N Ok 5 Ok Rate 1 Pass At 1 0.8 Pass At 1 P05 0.2 Pass At 1 P50 1 Pass At 1 P95 1 Timeout Rate 0 TTFT P50 28.0842 ms Total P50 Ms 4338.3245 Tokens Out Total 1653 Run configuration Model: meta-llama/Llama-3.1-70B-Instruct @ unknown00 Engine: vllm vunknown… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-70b-instruct__code-generation-humaneval-mini__019e3b935675.text-generationn<1K0 likes25 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.