Team Ai
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.1k downloads10mo agoHugging Face02wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes214 downloads2y agoHugging Face03thesven /CodeMaster-Phi-Instruct Code Master Phi is a compiled dataset designed for training Phi3 instruct models. This dataset is focused on code-based data and integrates multiple high-quality sources to ensure a robust training foundation. The sources include: Replete-AI/code_bagel: A diverse collection of code snippets and examples. nickrosh/Evol-Instruct-Code-80k-v1: A dataset featuring evolved instructions for code generation tasks. iamtarun/python_code_instructions_18k_alpaca: A compilation of Python code… See the full description on the dataset page: https://huggingface.co/datasets/thesven/CodeMaster-Phi-Instruct.texttext-generation1M<n<10M0 likes179 downloads2y agoHugging Face04liodon-ai /gemma4-code-review-instruct gemma4-code-review-instruct 197K code review examples — 58K with chain-of-thought <think> reasoning traces. Built to train models that don't just flag issues, but explain their reasoning before delivering a review. Drop-in ready for SFT with any chat model. Why This Dataset Most code review datasets give you diff → comment. This one gives you diff → think → comment for 30% of examples — reasoning traces that show how to analyze a diff before writing the review.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/gemma4-code-review-instruct.texttext-generation100K<n<1M4 likes133 downloads4mo agoHugging Face05JulianAT /SynthUI-Code-Instruct-2k-v1Synth UI 🎹 https://www.synthui.design Dataset details This dataset aims to provide a diverse collection of NextJS code snippets, along with their corresponding instructions, to facilitate the training of language models for NextJS-related tasks. It is designed to cover a wide range of NextJS functionalities, including UI components, routing, state management, and more. This dataset consists of: Note: The dataset is seperated into two main parts: raw Contains only the… See the full description on the dataset page: https://huggingface.co/datasets/JulianAT/SynthUI-Code-Instruct-2k-v1.texttext-generation1K<n<10K0 likes93 downloads2y agoHugging Face06kd13 /CodeChat-Instruct-v1 CodeChat-Instruct-v1 CodeChat-Instruct-v1 is a synthetic coding instruction dataset designed for supervised fine-tuning of language models on programming-related conversations. It includes diverse coding tasks such as code review, code improvement, complexity analysis, edge-case discussion, code explanation, library/API usage, refactoring guidance. The dataset is suitable for training coding assistants, educational programming tutors, and general-purpose code LLMs with strong… See the full description on the dataset page: https://huggingface.co/datasets/kd13/CodeChat-Instruct-v1.texttext-generation10K<n<100K1 likes78 downloads4mo agoHugging Face07rodriguescarson /adaption-code-oss-instruct-raw-aug OSS-Instruct Coding Tasks (Augmented) Coding problems inspired by open-source snippets, with solutions across several languages. Rows 8,996 Domain programming Format data.parquet, one row per example Licence other Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw-aug.texttext-generation1K<n<10K0 likes60 downloads14d agoHugging Face08rodriguescarson /adaption-code-oss-instruct-raw OSS-Instruct Coding Tasks Coding problems inspired by open-source snippets, with solutions across several languages. Rows 3,000 Domain programming Format data.parquet, one row per example Licence mit Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded. enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw.tabulartext-generation1K<n<10K0 likes58 downloads14d agoHugging Face09rodriguescarson /adaption-code-oss-instruct-raw-aug-e5bca4 OSS-Instruct Coding Tasks (Augmented) Coding problems inspired by open-source snippets, with solutions across several languages. Rows 7,000 Domain programming Format data.parquet, one row per example Licence other Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-code-oss-instruct-raw-aug-e5bca4.texttext-generation1K<n<10K0 likes58 downloads14d agoHugging Face10erythropygia /Instruct-Python-Code-Turkish Dataset Card for Instruct-Python-Code-Turkish Language: Turkish Dataset Description The translation was performed using the Google translation model to ensure high-quality, accurate translation. Dataset Details Size: ≈5K Translation tool: Google Translate Data format: Instruct, Output texttext-generation1K<n<10K1 likes55 downloads2y agoHugging Face11PotatoHD /code-instruct-mixed Description Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence. Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is. Usage from datasets import load_dataset ds = load_dataset("PotatoHD/code-instruct-mixed") texttext-generation100K<n<1M0 likes47 downloads4mo agoHugging Face12syntaxsynth /instruct_code_cleaning SFT code dataset building Contain a list of tasks useful when building a iniitial dataset source: reverse_translation Given a history of conversations, what would the human ask next? reverse_translation_first_round Suppose you already have a response, the LLM must predict what question does the human asked clean_code Given a code snippet, it determines whether its useful and atomic enough to be use for a response by LLM gen_code_question Generates a question given a… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/instruct_code_cleaning.texttext-generation10K<n<100K1 likes37 downloads3y agoHugging Face13jtatman /python-github-code-instruct-filtered-5k Dataset Card for "python-github-code-instruct-filtered-5k" This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03. Feedback and additional columns generated through OpenAI and Cohere responses. texttext-generation1K<n<10K7 likes35 downloads2y agoHugging Face14nayohan /Evol-Instruct-Code-80k-v1-koTranslated nickrosh/Evol-Instruct-Code-80k-v1 using nayohan/llama3-instrucTrans-enko-8b. This is a raw translation dataset. It needs to be filtered for repetitions generated by the model. texttext-generation10K<n<100K1 likes33 downloads2y agoHugging Face15msw-ai-tf /maplestory-worlds-creator-code-instruct MapleStory Worlds Creator Code (mlua) Instruction-style code dataset for mlua, the scripting language of MapleStory Worlds. Built from the official Creator Center example code: each example is grounded in its source document and paired with a natural-language task, reasoning, a self-contained explanation, and commented mlua code. Intended to teach LLMs to write mlua game scripts. The example code is preserved from the official source (a code-preservation check rejects any record… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-code-instruct.texttext-generation1K<n<10K2 likes24 downloads3mo agoHugging Face16tbilisi-ai-lab /code-instruct-ka code-instruct-ka Georgian code instruction dataset for training coding capabilities. Dataset Summary Property Value Examples 61,288 Size 937 MB Language Georgian Data Fields id: Unique identifier conversation: Conversation with code instructions Usage from datasets import load_dataset ds = load_dataset("tbilisi-ai-lab/code-instruct-ka") Citation @misc{tbilisi2025codeinstructka, title = {code-instruct-ka:… See the full description on the dataset page: https://huggingface.co/datasets/tbilisi-ai-lab/code-instruct-ka.texttext-generation10K<n<100K0 likes17 downloads5mo agoHugging Face17HachiML /amenokaku-code-instruct-python-mitkunishou/amenokaku-code-instructを以下の条件で絞り込んだものです。 MITライセンス (licence: 'MIT') python (source: ['gasyori_100_knocks', 'datascience_100_knocks_python', 'bifi', 'python_for_begginers_solve_50_exercises', 'nlp_100_knocks') texttext-generation1K<n<10K0 likes16 downloads2y agoHugging Face18HachiML /amenokaku-code-instruct-python-mit-450kunishou/amenokaku-code-instructを以下の条件で絞り込んだものです。 MITライセンス (licence: 'MIT') python (source: ['gasyori_100_knocks', 'datascience_100_knocks_python', 'bifi', 'python_for_begginers_solve_50_exercises', 'nlp_100_knocks') source: 'bifi'をランダムに100件に絞り込み texttext-generationn<1K0 likes15 downloads2y agoHugging Face19domofon /evol-instruct-code-cot-80k evol-instruct-code-cot-80k COT distilled dataset with 63,007 examples. Source Base: nickrosh/Evol-Instruct-Code-80k-v1 Model: Mistral-7B-Instruct-v0.2-AWQ Format instruction: Task thinking: <think>...</think> reasoning response: Solution texttext-generation10K<n<100K0 likes14 downloads10mo agoHugging Face20Sharathhebbar24 /Evol-Instruct-Code-80k-v1 Evol-Instruct-Code-80k-v1 This is a cleansed version of nickrosh/Evol-Instruct-Code-80k-v1 Usage from datasets import load_dataset dataset = load_dataset("Sharathhebbar24/Evol-Instruct-Code-80k-v1", split="train") texttext-generation10K<n<100K1 likes12 downloads3y agoHugging Face21kd13 /CodeDebug-Instruct-v2-Reasoning CodeDebug-Instruct-v2-Reasoning CodeDebug-Instruct-v2-Reasoning is a synthetic debugging instruction dataset designed for supervised fine-tuning of language models with enhanced code reasoning capabilities. It covers diverse debugging scenarios including import errors, syntax errors, runtime errors, performance bottlenecks, time limit exceeded (TLE) issues, and general code repair tasks with step-by-step reasoning and corrected solutions. The dataset is suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/kd13/CodeDebug-Instruct-v2-Reasoning.texttext-generation1K<n<10K0 likes12 downloads3mo agoHugging Face221-800-SHARED-TASKS /code_instruct_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K1 likes11 downloads2y agoHugging Face23kd13 /CodeDebug-Instruct-v1 CodeDebug-Instruct-v1 CodeDebug-Instruct-v1 is a synthetic debugging instruction dataset designed for supervised fine-tuning of language models on code error diagnosis and problem solving. It covers a wide range of debugging scenarios including import errors, syntax errors, runtime errors, performance bottlenecks, time limit exceeded (TLE) issues, and general code fixing tasks with clear explanations and corrected solutions. The dataset is suitable for training coding assistants… See the full description on the dataset page: https://huggingface.co/datasets/kd13/CodeDebug-Instruct-v1.texttext-generation10K<n<100K0 likes11 downloads4mo agoHugging Face24JuDDGES /en-appealcourt-coded-instruct_v02 Dataset Card for JuDDGES/en-appealcourt-coded-instruct_v02 Dataset Summary The raw data was acquired from publicly available judgments from the Court of Appeal (Criminal Division) (link) of England and Wales. These judgments are available in HTML format on the national archives website for online reading. They can be downloaded as XML or PDF files under the crown copyright license (link): and the Open Government license (see Appendix 6 in the paper). These licenses… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/en-appealcourt-coded-instruct_v02.texttext-generationn<1K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.