Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K173 likes47k downloads3y agoHugging Face02ise-uiuc /Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process). texttext-generation100K<n<1M190 likes30k downloads3y agoHugging Face03m-a-p /CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement [🏠Homepage] | [🛠️Code] OpenCodeInterpreter OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities. For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.textquestion-answering100K<n<1M211 likes24k downloads3y agoHugging Face04likaixin /InstructCoder Paper | Code | Blog InstructCoder (CodeInstruct): Empowering Language Models to Edit Code Updates May 23, 2023: Paper, code and data released. Overview InstructCoder is the first dataset designed to adapt LLMs for general code editing. It consists of 114,239 instruction-input-output triplets and covers multiple distinct code editing scenarios, generated by ChatGPT. LLaMA-33B finetuned on InstructCoder performs on par with ChatGPT on a… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/InstructCoder.texttext-generation100K<n<1M17 likes11k downloads2y agoHugging Face05nickrosh /Evol-Instruct-Code-80k-v1Open Source Implementation of Evol-Instruct-Code as described in the WizardCoder Paper. Code for the intruction generation can be found on Github as Evol-Teacher. text10K<n<100K251 likes7.7k downloads3y agoHugging Face06nvidia /Nemotron-SFT-Instruction-Following-Chat-v3 Dataset Description: The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following. The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.texttext-generation100K<n<1M20 likes6k downloads4mo agoHugging Face07wis-k /instruction-following-evaltextn<1K10 likes5k downloads3y agoHugging Face08Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K135 likes5k downloads1y agoHugging Face09ICTNLP /InstructS2S-200K InstructS2S-200K Dataset Description InstructS2S-200K is a multi-turn speech-to-speech conversation dataset containing approximately 200,000 dialogues, developed for the LLaMA-Omni and LLaMA-Omni 2 research projects on real-time spoken chatbots. Usage The dataset is split into multiple parts and needs to be reconstructed: # Combine the parts and extract cat en_part_* > instructs2s_200k.tar.gz tar -xzf instructs2s_200k.tar.gz License This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/InstructS2S-200K.text100K<n<1M11 likes3.9k downloads10mo agoHugging Face10TokenBender /code_instructions_122k_alpaca_styletext100K<n<1M80 likes3.3k downloads3y agoHugging Face11mesolitica /Malay-Dialect-Instructions Malay dialect instruction including coding Negeri Sembilan QA public transport QA, Coding CUDA coding, Kedah QA infra QA, Coding Rust coding, Kelantan QA Najib Razak QA, Coding Go coding, Perak QA Anwar Ibrahim QA, Coding SQL coding, Pahang QA Pendatang asing QA, Coding Typescript coding, Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.texttext-generation10K<n<100K6 likes3.2k downloads2y agoHugging Face12WizardLMTeam /WizardLM_evol_instruct_V2_196k News 🔥 🔥 🔥 [08/11/2023] We release WizardMath Models. 🔥 Our WizardMath-70B-V1.0 model slightly outperforms some closed-source LLMs on the GSM8K, including ChatGPT 3.5, Claude Instant 1 and PaLM 2 540B. 🔥 Our WizardMath-70B-V1.0 model achieves 81.6 pass@1 on the GSM8k Benchmarks, which is 24.8 points higher than the SOTA open-source LLM. 🔥 Our WizardMath-70B-V1.0 model achieves 22.7 pass@1 on the MATH Benchmarks, which is 9.2 points higher than the SOTA open-source LLM.… See the full description on the dataset page: https://huggingface.co/datasets/WizardLMTeam/WizardLM_evol_instruct_V2_196k.text100K<n<1M252 likes2.5k downloads3y agoHugging Face13nvidia /Nemotron-Instruction-Following-Chat-v1 Dataset Description: The Nemotron-Instruction-Following-Chat-v1 dataset is designed to broadly strengthen the model’s interactive capabilities, spanning open-ended chat, precise instruction following, and reliable structured output generation. It combines refreshed chat data from Nemotron-Post-Training-Dataset-v2 (extended to multi-turn) with synthetic dialogues produced by strong frontier models such as GPT-OSS-120B and Qwen3-235B variants. This dataset is ready for commercial… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Instruction-Following-Chat-v1.text100K<n<1M132 likes2.5k downloads10mo agoHugging Face14MemGPT /MSC-Self-Instruct MemGPT This is the self-instruct dataset of MSC conversations used for MemGPT paper. For more information please refer to memgpt.ai The MSC dataset is a multi-round human conversations. In this dataset, our goal is to come up with a conversation opener, that is personalized to the user by referencing topics from the previous conversations. These were generated while evaluating MemGPT. textn<1K13 likes2.4k downloads3y agoHugging Face15philschmid /trl-test-instructiontextn<1K0 likes2.2k downloads3y agoHugging Face16gussieIsASuccessfulWarlock /security_instruct_mcq_2481textn<1K0 likes2.2k downloads2y agoHugging Face17Vividbot /vivid-video-instructtext10K<n<100K2 likes2.1k downloads2y agoHugging Face18bcb-instruct /bcb_datatabular10M<n<100M0 likes1.6k downloads1y agoHugging Face19OLMo-Coding /starcoder-python-instruct StarCoder-Python-Qwen-Instruct Dataset Description This dataset contains Python code samples paired with synthetically generated natural language instructions. It is designed for supervised fine-tuning of language models for code generation tasks. The dataset is derived from the Python subset of the bigcode/starcoderdata corpus, and the instructional text for each code sample was generated using the Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 model. Creation… See the full description on the dataset page: https://huggingface.co/datasets/OLMo-Coding/starcoder-python-instruct.text1M<n<10M14 likes1.3k downloads1y agoHugging Face20anthracite-org /kalo-opus-instruct-22k-no-refusaltext10K<n<100K39 likes1.1k downloads2y agoHugging Face21WizardLMTeam /WizardLM_evol_instruct_70kThis is the training data of WizardLM. News 🔥 🔥 🔥 [08/11/2023] We release WizardMath Models. 🔥 Our WizardMath-70B-V1.0 model slightly outperforms some closed-source LLMs on the GSM8K, including ChatGPT 3.5, Claude Instant 1 and PaLM 2 540B. 🔥 Our WizardMath-70B-V1.0 model achieves 81.6 pass@1 on the GSM8k Benchmarks, which is 24.8 points higher than the SOTA open-source LLM. 🔥 Our WizardMath-70B-V1.0 model achieves 22.7 pass@1 on the MATH Benchmarks, which is 9.2 points… See the full description on the dataset page: https://huggingface.co/datasets/WizardLMTeam/WizardLM_evol_instruct_70k.text10K<n<100K199 likes1.1k downloads3y agoHugging Face22re-align /just-eval-instruct Just Eval Instruct Highlights Data sources: AlpacaEval (covering 5 datasets), LIMA-test, MT-bench, Anthropic red-teaming, and MaliciousInstruct. 1K examples: 1,000 instructions, including 800 for problem-solving test, and 200 specifically for safety test. Category: We tag each example with (one or multiple) labels on its task types and topics.… See the full description on the dataset page: https://huggingface.co/datasets/re-align/just-eval-instruct.text10K<n<100K34 likes991 downloads3y agoHugging Face23nvidia /Nemotron-RL-Instruction-Following-Calendar-v2 Dataset Description: The Calendar-Scheduling-Dataset is a multi-turn conversation dataset that can understand natural language scheduling constraints, follow instructions across multiple messages, infer scheduling conflicts and satisfy multiple constraints simultaneously. Each event has constraints around duration (e.g. 45 min) and timing (e.g. should be scheduled after 3pm). The user mentions the events and associated constraints in a random order in a natural conversational… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Calendar-v2.text1K<n<10K5 likes967 downloads12d agoHugging Face24RUC-DataLab /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/DataScience-Instruct-500K.tabular10K<n<100K76 likes965 downloads1y agoHugging Face25Multilingual-Multimodal-NLP /IfEvalCode-Instructtext1K<n<10K2 likes962 downloads1y agoHugging Face26Mohammed-Altaf /medical-instruction-120k What is the Dataset About?🤷🏼‍♂️ The dataset is useful for training a Generative Language Model for the Medical application and instruction purposes, the dataset consists of various thoughs proposed by the people [mentioned as the Human ] and there responses including Medical Terminologies not limited to but including names of the drugs, prescriptions, yogic exercise suggessions, breathing exercise suggessions and few natural home made prescriptions. How the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mohammed-Altaf/medical-instruction-120k.text100K<n<1M11 likes876 downloads3y agoHugging Face27xzuyn /open-instruct-uncensored-alpacaOriginal dataset page from ehartford. 810,102 entries. Sourced from open-instruct-uncensored.jsonl. Converted the jsonl to a json which can be loaded into something like LLaMa-LoRA-Tuner. I've also included smaller datasets that includes less entries depending on how much memory you have to work with. Each one is randomized before being converted, so each dataset is unique in order. Count of each Dataset: code_alpaca: 19991 unnatural_instructions: 68231 baize: 166096 self_instruct: 81512… See the full description on the dataset page: https://huggingface.co/datasets/xzuyn/open-instruct-uncensored-alpaca.text1M<n<10M7 likes818 downloads3y agoHugging Face28nothingiisreal /Claude-3-Opus-Instruct-15K Original Character Card Processed 15K Prompts - See Usable Responses Below Based on Claude 3 Opus through AWS. I took a random 5K + 10K prompt subset from Norquinal/claude_multi_instruct_30k to use as prompts, and called API for my answers. Warning! Uncleaned - Only Filtered for Blatant Refusals. I will be going through and re-prompting missing prompts, but I do not expect much success, as some of the prompts shown are nonsensical, incomplete, or impossible… See the full description on the dataset page: https://huggingface.co/datasets/nothingiisreal/Claude-3-Opus-Instruct-15K.text10K<n<100K21 likes750 downloads2y agoHugging Face29bcb-instruct /datatabular1M<n<10M0 likes733 downloads1y agoHugging Face30mesolitica /instructions-pair-miningtext100K<n<1M2 likes730 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.