Team Ai
Datasetpublic

ikergf/guiasalud

GuiaSalud: structured clinical dataset for evidence-grounded open-answer QA GuiaSalud is an open-answer medical question-answering dataset built from six Spanish clinical practice guidebooks published by GuíaSalud (an organism that belongs to the Spanish Ministry of Health / Ministerio de Sanidad). It transforms unstructured clinical guideline text into 919 structured question-answer-evidence instances, each separating the question asked, the guideline panel's judgement, and the… See the full description on the dataset page: https://huggingface.co/datasets/ikergf/guiasalud.

sourceHugging Facecc-by-nc-4.0updated 10d agoView on Hugging Face
0likes134downloads
Dataset Card

GuiaSalud: structured clinical dataset for evidence-grounded open-answer QA

GuiaSalud is an open-answer medical question-answering dataset built from six Spanish clinical practice guidebooks published by GuíaSalud (an organism that belongs to the Spanish Ministry of Health / Ministerio de Sanidad). It transforms unstructured clinical guideline text into 919 structured question-answer-evidence instances, each separating the question asked, the guideline panel's judgement, and the evidence passage from the guideline that supports it.

The code used to construct, curate, validate, and export the dataset is publicly available in the GuiaSalud GitHub repository: https://github.com/iker-gutierrez/guiasalud.

The dataset is provided in two languages:

  • —`es/`: the original Spanish instances.
  • —`eu/`: a Basque translation of the same instances (machine-translated from the Spanish originals, with automatic checks for truncation and degenerate repetition applied to the translation pipeline).

Both language versions share the same instance IDs and schema, so a given id refers to the same underlying question in either language. Load the language you need as a separate configuration:

python
from datasets import load_dataset

spanish = load_dataset("ikergf/guiasalud", "es")
basque = load_dataset("ikergf/guiasalud", "eu")

Dataset structure

Each split is a JSON Lines file (train.jsonl, dev.jsonl, test.jsonl) under the corresponding language directory (es/ or eu/). Each line is one instance:

json
{
  "id": "guiasalud_1",
  "split": "train",
  "guidebook": "ansiedad.txt",
  "topic": "...",
  "subtopic": "...",
  "question": "...",
  "focus": "...",
  "judgement": "...",
  "evidence": "...",
  "considerations": "..."
}
  • —id: a globally unique instance identifier, shared across the es and eu versions.
  • —split: one of train, dev, or test.
  • —guidebook: the source clinical guideline file the instance was extracted from.
  • —topic: the clinical question the guideline recommendation addresses.
  • —subtopic: an optional refinement of the clinical topic.
  • —question: the specific sub-question asked (e.g. certainty of evidence, cost-effectiveness).
  • —focus: an optional refinement of the specific question.
  • —judgement: the guideline panel's short, direct answer for this sub-question.
  • —evidence: the guideline passage that supports the answer.
  • —considerations: additional considerations the guideline panel recorded alongside the evidence.

Splits

SplitInstances
train731
dev63
test125
Total919

Intended use and limitations

GuiaSalud is designed to support retrieval-augmented generation (RAG) and open-answer medical QA research in Spanish and Basque. It is used in MeviRAG, a RAG system for medical QA in both languages. The MeviRAG code is publicly available in the MeviRAG GitHub repository: https://github.com/iker-gutierrez/mevirag.

Basque is a low-resource language, and the Basque configuration is machine-translated and has not undergone comprehensive expert evaluation of linguistic and clinical accuracy. The dataset is intended for research and is not a clinically validated decision-support resource. It is derived from public clinical practice guidebooks and contains no patient records or personally identifiable information. Outputs from systems trained or evaluated with this dataset should not be treated as medical advice.

License

This dataset is released under the CC BY-NC 4.0 license.

Citation

If you use this dataset, please cite the Master’s thesis that introduces GuiaSalud and describes its construction, validation, and use in MeviRAG:

bibtex
@mastersthesis{gutierrezfandino2026mevirag,
  author = {Gutierrez Fandiño, Iker},
  title  = {{GuiaSalud dataset and MeviRAG: Towards evidence-grounded medical QA in Spanish and Basque}},
  school = {University of the Basque Country (EHU)},
  year   = {2026},
  type   = {Master's thesis}
}

Note: The Master's thesis was uploaded to ADDI (https://addi.ehu.eus) and will soon be openly available there.