Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K173 likes47k downloads3y agoHugging Face02ise-uiuc /Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process). texttext-generation100K<n<1M190 likes30k downloads3y agoHugging Face03iamtarun /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K351 likes20k downloads3y agoHugging Face04allenai /tulu-3-sft-personas-instruction-following Dataset Descriptions This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset. To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper. Curated by: Allen Institute for AI Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.texttext-generation10K<n<100K69 likes14k downloads2y agoHugging Face05likaixin /InstructCoder Paper | Code | Blog InstructCoder (CodeInstruct): Empowering Language Models to Edit Code Updates May 23, 2023: Paper, code and data released. Overview InstructCoder is the first dataset designed to adapt LLMs for general code editing. It consists of 114,239 instruction-input-output triplets and covers multiple distinct code editing scenarios, generated by ChatGPT. LLaMA-33B finetuned on InstructCoder performs on par with ChatGPT on a… See the full description on the dataset page: https://huggingface.co/datasets/likaixin/InstructCoder.texttext-generation100K<n<1M17 likes11k downloads2y agoHugging Face06nvidia /Nemotron-SFT-Instruction-Following-Chat-v3 Dataset Description: The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following. The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.texttext-generation100K<n<1M20 likes6k downloads4mo agoHugging Face07Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K135 likes5k downloads1y agoHugging Face08mesolitica /Malay-Dialect-Instructions Malay dialect instruction including coding Negeri Sembilan QA public transport QA, Coding CUDA coding, Kedah QA infra QA, Coding Rust coding, Kelantan QA Najib Razak QA, Coding Go coding, Perak QA Anwar Ibrahim QA, Coding SQL coding, Pahang QA Pendatang asing QA, Coding Typescript coding, Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.texttext-generation10K<n<100K6 likes3.2k downloads2y agoHugging Face09tasksource /tasksource-instruct tasksource-instruct Instruction-tuning data recast from the ~480 English classification, multiple-choice and token-classification tasks of tasksource. Every example comes from a human-built dataset (NLI, logical reasoning, sentiment, hate speech, discourse, argumentation, ...), not from a teacher model. Each task is capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2, for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct.texttext-generation1M<n<10M24 likes2.6k downloads13d agoHugging Face10nicholasKluge /Pt-Corpus-Instruct Portuguese-Corpus Instruct Dataset Summary Portuguese-Corpus Instruct is a concatenation of several portions of Brazilian Portuguese datasets found in the Hub. In a tokenized format, the dataset (uncompressed) weighs 80 GB and has approximately 6.2B tokens. This version of the corpus (Pt-Corpus-Instruct) includes several instances of conversational and general instructional data, allowing trained models to go through preference pre-training during their initial… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct.texttext-generation10M<n<100M3 likes1.9k downloads2y agoHugging Face11manifoldlabs /Infinity-Instruct Infinity Instruct Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] (would be released soon) The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and… See the full description on the dataset page: https://huggingface.co/datasets/manifoldlabs/Infinity-Instruct.texttext-generation10M<n<100M5 likes1.9k downloads2y agoHugging Face12BAAI /Infinity-Instructgated Infinity Instruct Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and high-quality instruction… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Infinity-Instruct.tabulartext-generation10M<n<100M765 likes1.9k downloads10mo agoHugging Face1304RR /tiny-instruct tiny-instruct-v1 This dataset is collated from multiple other open-source datasets (de-duplicated). This has a total of ~6M rows each with an instruction and response (single-turn converstion). Code Datasets: CodeAlpaca_20K CodeExercise-Python-27k Evol-Instruct-Code-80k-v1 tiny-codes Evol-instruction-66k sciphi-python-textbook programming_books_llama WizardLM_evol_instruct_70k Math Datasets: MetaMathQA arxiv-math-instruct-50k MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/04RR/tiny-instruct.texttext-generation1M<n<10M16 likes1.8k downloads3y agoHugging Face14solanaclawd /solana-clawd-realtime-research-instruct Solana Clawd Realtime Research Instruct 83,662 English instruction conversations for Solana mechanics, agent tools, research retrieval, protocol reasoning, and risk-aware analysis. This is a published data release from Solana Clawd — The Sovereign Agent Stack on Solana. The initiative builds ecosystem-native models that understand accounts, PDAs, versioned transactions, address lookup tables, Pump.fun graduation, and RPC failure modes. Its intended architecture separates the… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-realtime-research-instruct.texttext-generation10K<n<100K0 likes1.4k downloads9d agoHugging Face15aisingapore /Instruction-Following-IFEvalgated SEA-IFEval SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. It is based on IFEval and was manually translated by native speakers for Indonesian, Javanese, Sundanese, Thai, Tagalog, and Vietnamese. Supported Tasks and Leaderboards SEA-IFEval is designed for evaluating chat or instruction-tuned large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Instruction-Following-IFEval.texttext-generation1K<n<10K0 likes1.4k downloads9mo agoHugging Face16openeurollm /Dolci-Instruct-SFT-translatedtexttext-generation1M<n<10M3 likes1.2k downloads4mo agoHugging Face17danish-foundation-models /faroese-dyna-instruct 🧨 Faroese dyna-instruct Version 0.1.1 (Changelog) Language Faroese (fao) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 9.18K Number of tokens (Llama 3): 2.68M Average conversation length in tokens (min, max): 291.83 (27, 1.24K) Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.texttext-generation10K<n<100K2 likes1.2k downloads5d agoHugging Face18BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.1k downloads10mo agoHugging Face19nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K8 likes1k downloads12d agoHugging Face20matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes857 downloads3y agoHugging Face21sujet-ai /Sujet-Finance-Instruct-177k Sujet Finance Dataset Overview The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.tabulartext-generation100K<n<1M86 likes847 downloads3y agoHugging Face22causal-lm /instructions Merged Instructions Dataset Merged Dataset for the response of instructions. texttext-generation10M<n<100M26 likes836 downloads3y agoHugging Face23iamtarun /code_instructions_120k_alpaca Dataset Card for code_instructions_120k_alpaca This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here. texttext-generation100K<n<1M69 likes792 downloads3y agoHugging Face24andresnowak /Instruction-finetuning-mixture-mnlp-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed texttext-generation1M<n<10M0 likes757 downloads1y agoHugging Face25Mxode /Chinese-Instruct-Lite 中文指令微调数据集 - Lite 版本 💻 Github Repo [!TIP] 这不是 Chinese-Instruct 的子集,而是一个全新的简化数据集。 如果您想要一个可以真实使用、而不仅仅适用于学习的数据集,欢迎访问:Mxode/Chinese-Instruct 如果您想要一个更加简单易收敛、主题集中的数据集,可以访问:Mxode/I_Wonder_Why-Chinese 具体构成 本数据集包含如下 5 个子集,总数据量 10M+。 code:代码主题的指令数据集,数据量 1.2M+。 math:数学主题的指令数据集,数据量 1.7M+。 general:通用指令数据集,主题广泛,与 code 和 math 指令不重复,数据量 5.1M+。 math(reasoning):数学推理数据集,指令采样自 math 子集,可通过 id 关联,数据量 1.2M+。code(reasoning):代码推理数据集,指令采样自 code 子集,可通过 id 关联,数据量 700K+。 如何使用… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Instruct-Lite.textquestion-answering10M<n<100M14 likes748 downloads1y agoHugging Face26alespalla /chatbot_instruction_prompts Dataset Card for Chatbot Instruction Prompts Datasets Dataset Summary This dataset has been generated from the following ones: tatsu-lab/alpaca Dahoas/instruct-human-assistant-prompt allenai/prosocial-dialog The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model textquestion-answering100K<n<1M65 likes685 downloads2y agoHugging Face27llm-blender /mix-instruct MixInstruct Introduction This is the official realease of dataset MixInstruct for project LLM-Blender. This dataset contains 11 responses from the current popular instruction following-LLMs that includes: Stanford Alpaca FastChat Vicuna Dolly V2 StableLM Open Assistant Koala Baize Flan-T5 ChatGLM MOSS Moasic MPT We evaluate each response with auto metrics including BLEU, ROUGE, BERTScore, BARTScore. And provide pairwise comparison results by prompting ChatGPT for the… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/mix-instruct.texttext-generation100K<n<1M38 likes625 downloads3y agoHugging Face28lateesha-bhatia /sec-filings-qa-instruct SEC Filings Instruction-Tuning Dataset (Llama-3 Format) This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template. Dataset Details Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.textquestion-answering1K<n<10K0 likes621 downloads1mo agoHugging Face29Nan-Do /instructional_code-search-net-python Dataset Card for "instructional_code-search-net-python" Dataset Summary This is an instructional dataset for Python. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-python.texttext-generation100K<n<1M37 likes581 downloads3y agoHugging Face30axiong /pmc_llama_instructionsThis repo provides part of the dataset used for PMC-LLaMA-13B's instruction tuning. Data Size Link ChatDoctor 100K https://www.yunxiangli.top/ChatDoctor/ MedQA 10.2K https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options MedMCQA 183K https://huggingface.co/datasets/medmcqa PubmedQA 211K https://huggingface.co/datasets/pubmed_qa LiveQA 635 https://huggingface.co/datasets/truehealth/liveqa MedicationQA 690 https://huggingface.co/datasets/truehealth/medicationqa UMLS… See the full description on the dataset page: https://huggingface.co/datasets/axiong/pmc_llama_instructions.textquestion-answering100K<n<1M33 likes529 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.