Team Ai
Datasetpublic

snuh/essential-level_medical_knowledge_dataset_sft

essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.

sourceHugging Facecc-by-nd-4.0updated 19d agoView on Hugging Face
0likes80downloads
Dataset Card

essential-levelmedicalknowledgedatasetsft

Dataset Description

Dataset Summary

essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.

This dataset is derived and augmented from the [Essential Medical Knowledge Data (필수의료 의학지식 데이터)](https://aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&&srchDataRealmCode=REALM006&aihubDataSe=data&dataSetSn=71875) provided by AI-Hub.

It focuses exclusively on complex clinical scenarios generated using the "Add Constraints" method from Evol-Instruct. Unlike simple fact-retrieval datasets, these examples force the model to navigate conflicting clinical limitations (e.g., comorbidities, resource constraints, contraindications). Each example includes a detailed Reasoning Path generated by a high-performance thinking model (qwen3-235b-a22b-thinking-2507), ensuring the model learns the logical process behind the diagnosis and treatment plan.

Note: This repository contains only the augmented data (44,985 samples) derived from the original AI-Hub source data.

Supported Tasks

  • —Medical Question Answering: Answering complex clinical queries across various specialties.
  • —Clinical Reasoning: Generating step-by-step logical explanations (<think> tags) before providing the final answer.
  • —Domain Adaptation: Fine-tuning models for specific medical departments using the provided configs.

Dataset Structure

Data Instances

The dataset is provided in a JSONL format compatible with Chat templates. Each instance contains a conversation history where the assistant provides a hidden reasoning process within <think> tags before the final response.

json
{
  "messages": [
    {
      "role": "user",
      "content": "32세 여성은 임신 중기 이후 2회의 반복적 임신 손실 병력이 있으며... (중략) ... 이 환자에서 조기 분만 위험 감소를 위한 가장 적절한 치료 방침은 무엇인가?"
    },
    {
      "role": "assistant",
      "content": "<think>\n...(reasoning process)...\n</think>\n### **의료 문제 분석 및 단계별 추론 과정**\n...\n**정답: 3**..."
    }
  ]
}

Data Fields

  • —messages: A list of dictionaries representing the conversation.
  • —role: The role of the speaker (user or assistant).
  • —content: The text content.
  • —User Content: Contains the medical question with added constraints.
  • —Assistant Content: Contains the Thinking Process (<think>...</think>) followed by the Final Answer.

Dataset Statistics

This repository covers 4 Medical Domains. You can load the entire dataset or specific domains using the config_name.

Config NameDomain (KR)Domain (EN)Count (Augmented)Source
ob_gyn산부인과Obstetrics & Gynecology5,939AI-Hub
pediatrics소아과Pediatrics / Neonatology7,229AI-Hub
emergency응급의학과Emergency Medicine1,933AI-Hub
internal_medicine내과Internal Medicine29,884AI-Hub
all전체 합계Total44,985

Usage

You can load the entire dataset or a specific domain using the config_name parameter.

python
from datasets import load_dataset

# 1. Load ALL data (Default)
dataset_all = load_dataset("snuh/essential-level_medical_knowledge_dataset_sft", "all")

# 2. Load Specific Domain (e.g., Obstetrics & Gynecology)
dataset_ob_gyn = load_dataset("snuh/essential-level_medical_knowledge_dataset_sft", "ob_gyn")

# 3. Load Another Domain (e.g., Internal Medicine)
dataset_im = load_dataset("snuh/essential-level_medical_knowledge_dataset_sft", "internal_medicine")

# Inspect an example
print(dataset_ob_gyn['train'][0]['messages'])

Dataset Creation

1. Source Data

The original question-answer pairs were sourced from the [Essential Medical Knowledge Data (필수의료 의학지식 데이터)](https://aihub.or.kr/aihubdata/data/view.do?currMenu=115&topMenu=100&&srchDataRealmCode=REALM006&aihubDataSe=data&dataSetSn=71875) on AI-Hub. This dataset was selected for its high-quality, expert-verified medical content covering essential clinical domains.

2. Data Augmentation: Evol-Instruct (Add Constraints)

To bridge the gap between textbook knowledge and complex clinical reality, we applied the Evol-Instruct methodology, focusing exclusively on the "Add Constraints" technique.

Original simple questions were rewritten to include specific limitations that force the model to weigh conflicting factors:

  • —Comorbidities: Adding conditions like heart failure, renal failure, or diabetes that complicate standard treatments.
  • —Medication History: Introducing drug interactions (e.g., anticoagulants, immunosuppressants).
  • —Environmental/Resource Constraints: Scenarios with limited equipment or lack of specialists.
  • —Special Populations: Pregnancy, pediatric patients, elderly patients with multiple comorbidities.

3. Reasoning Path Generation

We utilized `qwen3-235b-a22b-thinking-2507` to generate high-quality reasoning paths for the augmented data. The generated reasoning follows a structured approach:

  • —Constraint Satisfaction: Verifying if each option satisfies the specific constraints added during augmentation.
  • —Hierarchical Elimination: Logically ruling out distractors based on priority and contraindications.

Limitations & Disclaimer

  • —Synthetic Nature: The examples are synthetically generated to increase complexity. While they mimic real-world scenarios, they are artificial constructs.
  • —AI-Generated Reasoning: The reasoning paths are generated by an AI model. Although they are designed to be logical and accurate, they may contain hallucinations or inaccuracies.
  • —Not for Clinical Use: This dataset is intended for research and training purposes only. It should not be used as a substitute for professional medical advice, diagnosis, or treatment.

Citation

If you use this dataset, please cite the following:

bibtex
@misc{essential_level_medical_knowledge_dataset_sft,
    title     = {essential-level_medical_knowledge_dataset_sft},
    url       = {https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft},
    author    = {Healthcare AI Research Institute (HARI) of Seoul National University Hospital (SNUH)},
    month     = {January},
    year      = {2026}
}