Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K173 likes47k downloads3y agoHugging Face02ise-uiuc /Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process). texttext-generation100K<n<1M190 likes30k downloads3y agoHugging Face03likaixin /InstructCoder Paper | Code | Blog InstructCoder (CodeInstruct): Empowering Language Models to Edit Code Updates May 23, 2023: Paper, code and data released. Overview InstructCoder is the first dataset designed to adapt LLMs for general code editing. It consists of 114,239 instruction-input-output triplets and covers multiple distinct code editing scenarios, generated by ChatGPT. LLaMA-33B finetuned on InstructCoder performs on par with ChatGPT on a… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/InstructCoder.texttext-generation100K<n<1M17 likes11k downloads2y agoHugging Face04nvidia /Nemotron-SFT-Instruction-Following-Chat-v3 Dataset Description: The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following. The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.texttext-generation100K<n<1M20 likes6k downloads4mo agoHugging Face05Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K135 likes5k downloads1y agoHugging Face06mesolitica /Malay-Dialect-Instructions Malay dialect instruction including coding Negeri Sembilan QA public transport QA, Coding CUDA coding, Kedah QA infra QA, Coding Rust coding, Kelantan QA Najib Razak QA, Coding Go coding, Perak QA Anwar Ibrahim QA, Coding SQL coding, Pahang QA Pendatang asing QA, Coding Typescript coding, Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.texttext-generation10K<n<100K6 likes3.2k downloads2y agoHugging Face07llm-blender /mix-instruct MixInstruct Introduction This is the official realease of dataset MixInstruct for project LLM-Blender. This dataset contains 11 responses from the current popular instruction following-LLMs that includes: Stanford Alpaca FastChat Vicuna Dolly V2 StableLM Open Assistant Koala Baize Flan-T5 ChatGLM MOSS Moasic MPT We evaluate each response with auto metrics including BLEU, ROUGE, BERTScore, BARTScore. And provide pairwise comparison results by prompting ChatGPT for the… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/mix-instruct.texttext-generation100K<n<1M38 likes625 downloads3y agoHugging Face08axiong /pmc_llama_instructionsThis repo provides part of the dataset used for PMC-LLaMA-13B's instruction tuning. Data Size Link ChatDoctor 100K https://www.yunxiangli.top/ChatDoctor/ MedQA 10.2K https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options MedMCQA 183K https://huggingface.co/datasets/medmcqa PubmedQA 211K https://huggingface.co/datasets/pubmed_qa LiveQA 635 https://huggingface.co/datasets/truehealth/liveqa MedicationQA 690 https://huggingface.co/datasets/truehealth/medicationqa UMLS… See the full description on the dataset page: https://huggingface.co/datasets/axiong/pmc_llama_instructions.textquestion-answering100K<n<1M33 likes529 downloads3y agoHugging Face09Multilingual-Multimodal-NLP /McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval. texttext-generation10K<n<100K39 likes452 downloads2y agoHugging Face10FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes447 downloads1y agoHugging Face11proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes435 downloads6mo agoHugging Face12aarajbhattarai /law-instructions-dataset Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.texttext-generation1K<n<10K0 likes386 downloads27d agoHugging Face13Mxode /Chinese-Instruct 中文指令微调数据集 💻 Github Repo 本项目旨在构建一个高质量、多领域、大规模的中文指令微调数据集。 本项目将会持续更新。更多数据集欢迎访问 Github Repo。 [!TIP] 如果您想要一个可用于学习的简化版中文指令数据集,可以访问:Mxode/Chinese-Instruct-Lite 具体构成 dpsk-r1-distil:中文 DeepSeek-R1 蒸馏数据集,来自 Congliu/Chinese-DeepSeek-R1-Distill-data-110k,根据打分质量做了筛选,提取了最终的回答,未包含思考过程。 chinese-reasoning-distil:中文推理蒸馏数据集,来自 Mxode/Chinese-Reasoning-Distil-Data,提取了最终的回答,未包含思考过程。 firefly:中文通用指令微调数据集,指令取自 Mxode/Firefly-1.1M-Rephrased,其本身已经相较于原 Firefly… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Instruct.texttext-generation1M<n<10M150 likes385 downloads1y agoHugging Face14Frost2o24 /bash-instruct-III-55k Bash Instruct III — 54,360 verified natural-language → Bash pairs Bash Instruct III is a synthetic instruction-tuning dataset that maps natural-language requests to correct Bash: single commands, short pipelines, and multi-line scripts. It is built for supervised fine-tuning of small and mid-size LLMs that must turn a plain request into shell code that actually runs. Every row is a three-turn chat conversation (system / user / assistant) with metadata for slicing (category… See the full description on the dataset page: https://huggingface.co/datasets/Frost2o24/bash-instruct-III-55k.texttext-generation10K<n<100K0 likes350 downloads1mo agoHugging Face15ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes345 downloads3y agoHugging Face16nvidia /Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1 Dataset Description: Teaches the model to follow arbitrary text formatting instructions (bullet styles, numbering, delimiters, heading formats, inline emphasis, web-answer structure, etc.) for targeted chat behaviors. Uses explicit Regex and string matching for the reward signal. This dataset is ready for commercial or non-commercial uses. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: Created on: April 10, 2026 Last Modified on: April… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.texttext-generation1K<n<10K3 likes331 downloads12d agoHugging Face17FreedomIntelligence /RAG-Instruct Introduction RAG-Instruct is a RAG dataset designed to comprehensively enhance LLM RAG capabilities, synthesized using GPT-4o. This dataset is based on the Wikipedia corpus and This dataset is based on the Wikipedia corpus and offers the advantages of query-document scenario diversity and task diversity. The RAG-Instruct dataset can significantly enhance the RAG ability of LLMs and make remarkable improvements in RAG performance across various tasks. Model WQA (acc) PQA (acc)… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/RAG-Instruct.textquestion-answering10K<n<100K48 likes321 downloads2y agoHugging Face18silk-road /Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集 Wizard-LM包含了很多难度超过Alpaca的指令。 中文的问题翻译会有少量指令注入导致翻译失败的情况 中文回答是根据中文问题再进行问询得到的。 我们会陆续将更多数据集发布到hf,包括 Coco Caption的中文翻译 CoQA的中文翻译 CNewSum的Embedding数据 增广的开放QA数据 WizardLM的中文翻译 如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。 骆驼(Luotuo): 开源中文大语言模型 https://github.com/LC1332/Luotuo-Chinese-LLM 骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。 ( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 ) 骆驼项目不是商汤科技的官方产品。 Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.texttext-generation10K<n<100K98 likes320 downloads3y agoHugging Face19ctaxnagomi /INSTRUCT_JEV INSTRUCT_JEV INSTRUCT_JEV is an instruction corpus built from the TypeSafe AI documentation for Jev, the first System One model. It is structured around the three TypeSafe question primitives - Choice, Noul and Score - and mirrors the raw corpus captured in deckerGUI-jev_corpus_RAW. Credits INSTRUCT_JEV is a DeckerGUI project and exists because of the work below. Who Contribution Link TypeSafe AI Jev - the first System One model - and the Choice / Noul… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/INSTRUCT_JEV.tabulartext-generationn<1K3 likes317 downloads21d agoHugging Face20nvidia /Nemotron-RL-Instruction-Following-Citation-Formatting-v1 Dataset Description: Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations. This dataset is ready for commercial/non-commercial uses. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: Created on: April 10, 2026 Last Modified on: April 10, 2026 Version: Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.texttext-generation1K<n<10K3 likes313 downloads12d agoHugging Face21agentlans /rombodawg-Everything_Instruct Everything-Instruct: Supervised Finetuning Dataset This dataset contains over 7 000 000 instruction-response pairs for supervised fine-tuning large language models. It combines the following datasets: rombodawg/Everything_Instruct rombodawg/Everything_Instruct_Multilingual It can be used for: Improving code generation and debugging Enhancing creative writing Improving general instruction followingFor English and many other languages Processing Removing duplicate… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/rombodawg-Everything_Instruct.texttext-generation1M<n<10M0 likes310 downloads10mo agoHugging Face22turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M41 likes298 downloads2y agoHugging Face23NickIBrody /python-code-instructions-85k Python Code Instructions - 85K Instruction-tuning dataset of Python functions paired with short natural-language instructions derived from repository docstrings. What changed in this release This release keeps the original public rows and format, but makes the dataset easier to use responsibly: exact duplicate rows were removed again using normalized instruction + output hashing deterministic train, validation, and test splits were added the dataset card now documents… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/python-code-instructions-85k.texttext-generation10K<n<100K1 likes298 downloads6mo agoHugging Face24Lev384501 /russian_instructions_2_cleaned Russian Instructions Cleaned Очищенная версия Den4ikAI/russian_instructions_2. Что сделано Дедупликация по question (удалено ~45k) Удалены пустые question и answer Удалены ответы короче 100 и длиннее 4000 символов Конвертировано в chat-формат (messages: user/assistant) Статистика Метрика Значение Исходно 237 281 После чистки 138 973 Удалено 98 308 (41%) Формат JSONL, одна строка = один пример. {"messages":… See the full description on the dataset page: https://huggingface.co/datasets/Lev384501/russian_instructions_2_cleaned.texttext-generation100K<n<1M2 likes264 downloads23d agoHugging Face25ed001 /ds-coder-instruct-v2 Dataset Card for DS Coder Instruct v2 Dataset Changes from v1: Added WizardLM evol data science samples Removed R samples from v2 DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in Python (R samples were removed in v2). The goal of this dataset is to enable… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v2.tabulartext-generation10K<n<100K13 likes263 downloads3y agoHugging Face26schneiderkamplab /dfm13-multilingual-grounded-instruct-hu dfm13_wave4_synthetic_hu_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-hu.texttext-generation10K<n<100K0 likes254 downloads6d agoHugging Face27schneiderkamplab /dfm13-multilingual-grounded-instruct-bg dfm13_wave4_synthetic_bg_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-bg.texttext-generation10K<n<100K0 likes253 downloads6d agoHugging Face28WrittenWithRust /Magicoder-OSS-Instruct-Rust-cleaned-3.9K 🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned) Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects. This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format. ⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.texttext-generation1K<n<10K1 likes252 downloads1mo agoHugging Face29schneiderkamplab /dfm13-multilingual-grounded-instruct-sk dfm13_wave4_synthetic_sk_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-sk.texttext-generation10K<n<100K0 likes237 downloads6d agoHugging Face30schneiderkamplab /dfm13-multilingual-grounded-instruct-sl dfm13_wave4_synthetic_sl_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-sl.texttext-generation10K<n<100K0 likes231 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.