Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
0104RR /tiny-instruct tiny-instruct-v1 This dataset is collated from multiple other open-source datasets (de-duplicated). This has a total of ~6M rows each with an instruction and response (single-turn converstion). Code Datasets: CodeAlpaca_20K CodeExercise-Python-27k Evol-Instruct-Code-80k-v1 tiny-codes Evol-instruction-66k sciphi-python-textbook programming_books_llama WizardLM_evol_instruct_70k Math Datasets: MetaMathQA arxiv-math-instruct-50k MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/04RR/tiny-instruct.texttext-generation1M<n<10M16 likes1.8k downloads3y agoHugging Face02merve /turkish_instructionstext10K<n<100K65 likes868 downloads3y agoHugging Face03sujet-ai /Sujet-Finance-Instruct-177k Sujet Finance Dataset Overview The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.tabulartext-generation100K<n<1M86 likes847 downloads3y agoHugging Face04mikheevshow /SIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-Instructtextn<1K0 likes374 downloads1y agoHugging Face05mikheevshow /SIGNAL-Dataset-Hiddens-Qwen-Qwen3-4B-Instruct-FP8This dataset contains hidden states of Qwen3-4B-Instruct model generated using SIGNAL Dataset. Sentence tokenization from transformers import AutoTokenizer from datasets import load_dataset tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Instruct-2507") # TBD textn<1K0 likes309 downloads1y agoHugging Face06jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes171 downloads6mo agoHugging Face07alxfgh /ChEMBL_Drug_Instruction_Tuning Dataset Card for ChEMBL Drug Instruction Tuning Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.textquestion-answering100K<n<1M15 likes152 downloads3y agoHugging Face08harryxi /HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-dataset-train-generationstext1M<n<10M0 likes129 downloads1y agoHugging Face09jamesdborin /Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.tabular1K<n<10K0 likes124 downloads3mo agoHugging Face10FinLang /investopedia-instruction-tuning-dataset Dataset Card for investopedia-instruction-tuning dataset We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.text100K<n<1M24 likes120 downloads2y agoHugging Face11Shiveswarran /llm_instruction_code_manual_yolo_lctextn<1K5 likes119 downloads3y agoHugging Face12md-nishat-008 /Bangla-Instruct Accepted in ACL Main 2025 TigerLLM - A Family of Bangla Large Language Models Nishat Raihan, Marcos Zampieri George Mason University, VA, USA mraihan2@gmu.edu If you find our work helpful, please consider citing our paper: @inproceedings{raihan-zampieri-2025-tigerllm, title = "{T}iger{LLM} - A Family of {B}angla Large Language Models", author = "Raihan, Nishat and Zampieri, Marcos", editor = "Che, Wanxiang and Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.texttext-generation100K<n<1M8 likes114 downloads1y agoHugging Face13alxfgh /PubChem_Drug_Instruction_Tuningtext10K<n<100K11 likes103 downloads3y agoHugging Face14Shiveswarran /llm_instruction_code_V6.1text100K<n<1M7 likes98 downloads3y agoHugging Face15Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes96 downloads9mo agoHugging Face16crosslingual-em /Qwen2.5-7B-Instruct-em-evaldocumentn<1K0 likes90 downloads5mo agoHugging Face17AddisGPT /AddisGPT-Amharic-Instruction AddisGPT-Amharic-Instruction A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions. 796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.tabulartext-generationn<1K1 likes80 downloads1mo agoHugging Face18lime-nlp /safer-instruct Safer-Instruct: Aligning Language Models with Automated Preference Data This repository contains the dataset for the paper titled "Safer-Instruct: Aligning Language Models with Automated Preference Data". Check out our project website here! Abstract Reinforcement learning from human feedback (RLHF) is a vital strategy for enhancing model capability in language models. However, annotating preference data for RLHF is a resource-intensive and creativity-demanding process… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/safer-instruct.text10K<n<100K1 likes75 downloads2y agoHugging Face19mikheevshow /SIGNAL-Dataset-Hiddens-Qwen-Qwen2.5-7B-Instructtextn<1K0 likes68 downloads1y agoHugging Face20taskydata /Pile-T5-Instruction_updatedtext10K<n<100K0 likes65 downloads2y agoHugging Face21DataFog /medical-transcription-instruct About This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field Dataset Summary Source: Original medical transcriptions with added instruction-output pairs Size: 38,924 instruction-output pairs Format: CSV file Domain: Medical / Healthcare Language: English Last Updated: 08-20-2024 Dataset Structure Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.tabular10K<n<100K31 likes62 downloads2y agoHugging Face22skyylord /mitre-attack-ttp-labeled-instructions MITRE ATT&CK TTP Mapping Dataset Training and evaluation data for mapping adversarial behavior descriptions (CTI reports, CTF writeups, CISA advisories) to MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). Built as my individual contribution to a research project conducted at LORIA (supervised by Jean-Yves Marion). This dataset was developed and used to fine-tune skyylord/qwen3-emb-0.6b-ttp with CachedMultipleNegativesRankingLoss and ANCE-style hard negative re-mining.… See the full description on the dataset page: https://huggingface.co/datasets/skyylord/mitre-attack-ttp-labeled-instructions.texttext-classification10K<n<100K0 likes60 downloads20d agoHugging Face23jean1 /45k_python_code_chinese_instruction Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details 中文提示的代码数据集 其中提示部分通过调用GPT-4.0-turbo API翻译成中文 Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/jean1/45k_python_code_chinese_instruction.text10K<n<100K6 likes58 downloads2y agoHugging Face24Raftico /instructional-dialogues-multilingual Multilingual Instructional Dialogues (10-Language Dataset) Multilingual Instructional Dialogues is a high-quality dataset of 100 structured, goal-oriented dialogues in 10 major world languages, created for training and fine-tuning AI assistants, chatbots, and instruction-tuned large language models. Each dialogue simulates a clear, polite interaction where a user asks for guidance on how to perform a task, and the assistant responds with easy-to-follow steps. This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/Raftico/instructional-dialogues-multilingual.text1K<n<10K2 likes55 downloads1y agoHugging Face25groloch /stable_diffusion_prompts_instruct Stable diffusion prompts for instruction models fine-tuning Overview This dataset contains 80,000+ prompts summarized to make it easier to create instruction-tuned prompt enhancing models. Each row of the dataset contains two values: a short description of a image a full prompt corresponding to that description in a stable diffusion format Hope this dataset can help creating amazing apps ! How to use You can download and use the dataset easily using the… See the full description on the dataset page: https://huggingface.co/datasets/groloch/stable_diffusion_prompts_instruct.text10K<n<100K2 likes53 downloads2y agoHugging Face26MBZUAI /instructpoet-ar Arabic Poetry IFT Dataset Summary Arabic Poetry IFT is a large-scale instruction-following dataset for Arabic poetry understanding and co-creation. It supports four task families: generation, continuation, revision/restoration, and multiple-choice analysis. The dataset covers Modern Standard Arabic (MSA) and four regional Arabic varieties used in the instruction layer: Gulf, Levantine, Nile Valley, and North African Arabic. This release accompanies the ACL 2026 paper… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/instructpoet-ar.texttext-generationn<1K0 likes52 downloads6mo agoHugging Face27halilibr /collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources: Turkish Poems Cleaned Turkish Reading Comprehension Question Answering Dataset Stanford ALPaCA Cleaned Turkish Translated Turkish Poems Turkish Folk Song Lyrics The data has been merged and processed for quality and consistency to create this dataset. texttext-generation100K<n<1M11 likes49 downloads3y agoHugging Face28Shiveswarran /llm_instruction_code_v6text100K<n<1M4 likes48 downloads3y agoHugging Face29jamesdborin /Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only.tabular1K<n<10K0 likes47 downloads3mo agoHugging Face30taskydata /Pile-T5-Instructiontext10K<n<100K0 likes46 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.