datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-instruct
tiny-instruct-v1
This dataset is collated from multiple other open-source datasets (de-duplicated). This has a total of ~6M rows each with an instruction and response (single-turn converstion).
Code Datasets:
CodeAlpaca_20K
CodeExercise-Python-27k
Evol-Instruct-Code-80k-v1
tiny-codes
Evol-instruction-66k
sciphi-python-textbook
programming_books_llama
WizardLM_evol_instruct_70k
Math Datasets:
MetaMathQA
arxiv-math-instruct-50k
MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/04RR/tiny-instruct.turkish_instructionsSujet-Finance-Instruct-177k
Sujet Finance Dataset Overview
The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.SIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-InstructSIGNAL-Dataset-Hiddens-Qwen-Qwen3-4B-Instruct-FP8This dataset contains hidden states of Qwen3-4B-Instruct model generated using SIGNAL Dataset.
Sentence tokenization
from transformers import AutoTokenizer
from datasets import load_dataset
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Instruct-2507")
# TBD
Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.ChEMBL_Drug_Instruction_Tuning
Dataset Card for ChEMBL Drug Instruction Tuning
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-dataset-train-generationsNemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Calendar-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Calendar-v2-prompt-only.investopedia-instruction-tuning-dataset
Dataset Card for investopedia-instruction-tuning dataset
We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data
and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that
ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.llm_instruction_code_manual_yolo_lcBangla-Instruct
Accepted in ACL Main 2025
TigerLLM - A Family of Bangla Large Language Models
Nishat Raihan, Marcos Zampieri
George Mason University, VA, USA
mraihan2@gmu.edu
If you find our work helpful, please consider citing our paper:
@inproceedings{raihan-zampieri-2025-tigerllm,
title = "{T}iger{LLM} - A Family of {B}angla Large Language Models",
author = "Raihan, Nishat and
Zampieri, Marcos",
editor = "Che, Wanxiang and
Nabende, Joyce and… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Bangla-Instruct.PubChem_Drug_Instruction_Tuningllm_instruction_code_V6.1Hinglish_Dataset_instruction_and_rawQwen2.5-7B-Instruct-em-evalAddisGPT-Amharic-Instruction
AddisGPT-Amharic-Instruction
A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions.
796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.safer-instruct
Safer-Instruct: Aligning Language Models with Automated Preference Data
This repository contains the dataset for the paper titled "Safer-Instruct: Aligning Language Models with Automated Preference Data". Check out our project website here!
Abstract
Reinforcement learning from human feedback (RLHF) is a vital strategy for enhancing model capability in language models. However, annotating preference data for RLHF is a resource-intensive and creativity-demanding process… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/safer-instruct.SIGNAL-Dataset-Hiddens-Qwen-Qwen2.5-7B-InstructPile-T5-Instruction_updatedmedical-transcription-instruct
About
This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field
Dataset Summary
Source: Original medical transcriptions with added instruction-output pairs
Size: 38,924 instruction-output pairs
Format: CSV file
Domain: Medical / Healthcare
Language: English
Last Updated: 08-20-2024
Dataset Structure
Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.mitre-attack-ttp-labeled-instructions
MITRE ATT&CK TTP Mapping Dataset
Training and evaluation data for mapping adversarial behavior descriptions (CTI reports,
CTF writeups, CISA advisories) to MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs).
Built as my individual contribution to a research project conducted at LORIA (supervised by Jean-Yves Marion). This dataset was developed and used to fine-tune skyylord/qwen3-emb-0.6b-ttp with CachedMultipleNegativesRankingLoss and ANCE-style hard negative re-mining.… See the full description on the dataset page: https://huggingface.co/datasets/skyylord/mitre-attack-ttp-labeled-instructions.45k_python_code_chinese_instruction
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
中文提示的代码数据集
其中提示部分通过调用GPT-4.0-turbo API翻译成中文
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/jean1/45k_python_code_chinese_instruction.instructional-dialogues-multilingual
Multilingual Instructional Dialogues (10-Language Dataset)
Multilingual Instructional Dialogues is a high-quality dataset of 100 structured, goal-oriented dialogues in 10 major world languages, created for training and fine-tuning AI assistants, chatbots, and instruction-tuned large language models.
Each dialogue simulates a clear, polite interaction where a user asks for guidance on how to perform a task, and the assistant responds with easy-to-follow steps. This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/Raftico/instructional-dialogues-multilingual.stable_diffusion_prompts_instruct
Stable diffusion prompts for instruction models fine-tuning
Overview
This dataset contains 80,000+ prompts summarized to make it easier to create instruction-tuned prompt enhancing models. Each row of the dataset contains two values:
a short description of a image
a full prompt corresponding to that description in a stable diffusion format
Hope this dataset can help creating amazing apps !
How to use
You can download and use the dataset easily using the… See the full description on the dataset page: https://huggingface.co/datasets/groloch/stable_diffusion_prompts_instruct.instructpoet-ar
Arabic Poetry IFT
Dataset Summary
Arabic Poetry IFT is a large-scale instruction-following dataset for Arabic poetry understanding and co-creation. It supports four task families: generation, continuation, revision/restoration, and multiple-choice analysis. The dataset covers Modern Standard Arabic (MSA) and four regional Arabic varieties used in the instruction layer: Gulf, Levantine, Nile Valley, and North African Arabic.
This release accompanies the ACL 2026 paper… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/instructpoet-ar.collected-turkish-instructions-v0.1This dataset is the result of merging and cleaning data from the following sources:
Turkish Poems Cleaned
Turkish Reading Comprehension Question Answering Dataset
Stanford ALPaCA Cleaned Turkish Translated
Turkish Poems
Turkish Folk Song Lyrics
The data has been merged and processed for quality and consistency to create this dataset.
llm_instruction_code_v6Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1-prompt-only.Pile-T5-Instruction
