datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-symptom-triage-conversationalTRIAGE_Bench
TRIAGE-Bench: Testing Resolution of Inter-Authority Guideline Evidence
TRIAGE-Bench is a benchmark for evaluating how LLMs resolve conflicts between authoritative clinical knowledge sources. It covers 2,000 items across four conflict types, each with an explicit governing policy that defines policy-consistent correctness.
Key Features
2,000 benchmark items (500 per conflict type) grounded in real guideline and drug-label discrepancies
Four conflict types:… See the full description on the dataset page: https://huggingface.co/datasets/DarrenLoong/TRIAGE_Bench.roman-urdu-emergency-triage-benchmark
Can Small AI Models Handle Medical Questions Written in Roman Urdu?
Author: Waqar AliRepository Type: Evaluation Benchmark & Empirical Capability AuditModels Evaluated:
Qwen/Qwen2.5-3B-Instruct (2.82B parameters)
meta-llama/Llama-3.2-3B-Instruct (3.21B parameters)
google/gemma-2-2b-it (2.61B parameters)Artifacts Included: 30-case expert-annotated clinical evaluation suite, deterministic generation logs, and Colab execution notebook.
1. The blind spot I… See the full description on the dataset page: https://huggingface.co/datasets/waqarali5498/roman-urdu-emergency-triage-benchmark.jevlogs-log-triage-benchmark
Jev Logs log-triage benchmark
A labeled evaluation of Jev Logs on sanitized public logs. Jev Logs asks TypeSafe’s Jev, through Vercel AI Gateway, whether a log line is worth sending to an expensive reasoning model. This dataset is a public, token-accounted measurement of that routing decision, including the 0.3.0 in-memory cache and local retain rules.
This is not a production-log study. Labels come from Loghub. HDFS labels are block-level, then joined onto every line that… See the full description on the dataset page: https://huggingface.co/datasets/reachjalil/jevlogs-log-triage-benchmark.satellite-disruption-triage-aux-v2-1
Satellite Disruption Triage Aux v2.1
Self-contained real-image repair of ChrisRPL/satellite-disruption-triage-aux-v2.
This version keeps only resolvable BRIGHT real-image rows in VLM SFT files. Synthetic reasoning rows are separated, SEN12MSCR is excluded because the license is unknown, and xBD-Ukraine rows from v2 are excluded because their image references are not resolvable in the stated source repo.
Files
train_flat.jsonl / train_sft.jsonl: real-image train rows… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-disruption-triage-aux-v2-1.Multilingual_medical_symptom_triage
tags:
- medical
- healthcare
- classification
- outbreak-detection
- triage
- multilingual
- adaption
- india
Multilingual Medical Symptom Triage Dataset
Dataset Description
A Mutlilingual medical triage dataset containing 9,064 patient
symptom descriptions in Hindi, English, and Hinglish (code-mixed
Hindi-English), paired with triage recommendations and rich
clinical metadata. Designed for training multilingual triage
classification models and… See the full description on the dataset page: https://huggingface.co/datasets/Tulsiandhare/Multilingual_medical_symptom_triage.triagent
TriAgent: Multi-Agent Committee Predictions for Financial Sentiment
This dataset contains every per-sentence prediction behind TriAgent, a divergence-aware multi-agent routing framework for cost-efficient LLM inference.
Paper (CIKM 2026): doi.org/10.1145/3799682.3839978
arXiv: arxiv.org/abs/2607.19794
Code: github.com/graphuofm/TRIAGENT
Project page: graphuofm.github.io/TRIAGENT
Dataset summary
The dataset holds 25,607 rows across five configurations. For each… See the full description on the dataset page: https://huggingface.co/datasets/dingjiacheng/triagent.opensec-triage
opensec-triage 0.5.0
Synthetic English security alert data for disposition classification, counterfactual evaluation and small model training experiments.
Each example pairs a security observation with contextual evidence and an expected disposition. The task is to classify the supplied evidence rather than infer a disposition from the observable action alone.
Configurations
Configuration
Purpose
Splits
default
Main 50,000 row text classification… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/opensec-triage.chsa-triage-baseline-metricssharky-triage-states
🦈 sharky-triage-states
Training and evaluation states for sharky-0.5B (luispoveda93/sharky-0.5B) — a Jev/Kev-style decision model for pcap vulnerability triage. Renamed from minicpm4-pcap-triage-states (contents identical).
78,463 per-flow states rendered as hex bytes (not parsed prose), each with a 3-class verdict and a train/val/locked split tag.
Structure
Path
Contents
data/
v1 corpus: 49,745 states from rdpahalavan/UNSW-NB15 Packet-Bytes… See the full description on the dataset page: https://huggingface.co/datasets/luispoveda93/sharky-triage-states.triage-bench
TriageBench: Judging Intervention Priority
TriageBench provides LLM-council judgments of which steps to repair first in failed multi-agent executions. It extends Who&When with candidate rankings, individual judgments, short rationales, and ranking uncertainty, alongside the original decisive-error labels.
This is the dataset accompanying TRIAGE: Severity-Ranked Multi-Agent Failure Attribution, by Xizhi Wang, accepted to AACL 2026 (Main Conference). TRIAGE code is available on… See the full description on the dataset page: https://huggingface.co/datasets/jimmywang585/triage-bench.2026-09-14-colosseum-hospital-self-sacrificial-qwen36-difficult-advice-702-fixed-as-triage
colosseum_hospital self_sacrificial of LASR-Callum/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_hospital self_sacrificial of LASR-Callum/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch (mode=think), mixed-checkpoint team;… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-colosseum-hospital-self-sacrificial-qwen36-difficult-advice-702-fixed-as-triage.ms_marco_triage_ratedmedical-triage-500Medical Synthetic Triage Dataset — 500 Cases
This dataset contains 500 high-quality synthetic medical triage entries, generated through a structured rule-based methodology. It is designed for:
AI / LLM prototyping
Healthcare AI experiments
Medical triage classification research
Educational and academic use
Risk assessment modeling
RAG and prompt evaluation
Non-diagnostic medical AI training
Dataset Features
100% synthetic (no real patient data)
CSV and JSONL formats
Includes symptoms… See the full description on the dataset page: https://huggingface.co/datasets/syntech-ai/medical-triage-500.email-triage-action-seed
Email Triage Action Seed
A small, fully-synthetic seed dataset for fine-tuning small (3–5B) language models on action-oriented email triage — classifying an inbox message into a category, priority, and the actionable decision a triage assistant should take.
The schema goes beyond classification: it asks the model to choose what to do with each email, not just what bucket it falls into.
Schema
Every row is one labelled email with five core fields:
Field
Allowed… See the full description on the dataset page: https://huggingface.co/datasets/weijianzhg/email-triage-action-seed.vulnerability-triage
VULNERABILITY_TRIAGE
A preference dataset for VULNERABILITY_TRIAGE, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/vulnerability-triage.satellite-disruption-triage-aux-v1-3
satellite-disruption-triage-aux-v1-3
Civilian Conflict-Disruption Satellite VLM Dataset — Auxiliary / v1.3
This is an auxiliary dataset for training and evaluating Vision-Language Models (VLMs) to perform civilian conflict-disruption triage from paired satellite imagery. It is not a tactical intelligence dataset and not a canonical expert benchmark.
Scope & Purpose
The target task is detecting macro-scale civilian infrastructure disruption caused by war, armed conflict… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-disruption-triage-aux-v1-3.triageiq-dataset
TriageIQ — Customer Support Ticket Classification Dataset
Synthetic dataset of 2,000 customer support tickets, each labeled with three independent classification axes.
Schema
Each example is a JSON object:
{
"text": "I've been charged twice this month, please refund me ASAP.",
"sentiment": "negative", // positive | neutral | negative
"urgency": "high", // low | medium | high
"category": "billing" // billing | technical |… See the full description on the dataset page: https://huggingface.co/datasets/coldstart88/triageiq-dataset.vscode-bug-feature-triage
VS Code Bug vs Feature Request Triage
Dataset summary
1,993 prepared issue records from public microsoft/vscode issues, reduced to one binary task: classify the issue text as bug or feature-request. The splits are a frozen temporal holdout (80/10/10 by created_at within each class, seed 42) used by the GitHub Triage SLM Fine-Tuning Benchmark to compare fine-tuned small models against their base checkpoints on the same test set. Each record carries cleaned issue… See the full description on the dataset page: https://huggingface.co/datasets/Tilakoid/vscode-bug-feature-triage.latentsig-med-triage-router
LatentSig Medical Triage Router Dataset
1,000 verified medical triage tool-call samples — 500 English + 500 Hinglish — for fine-tuning Small Language Models (SLMs) as structured medical triage routers.
Overview
This dataset trains SLMs (1B–3B parameters) to act as reliable structured tool-callers for clinical medical triage. Given a patient symptom description, the model must:
Select the correct tool from 7 available medical tools
Output a valid JSON tool call… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/latentsig-med-triage-router.fintech-support-triage
Fintech Support Triage
A multi-turn customer-support dataset with a reward you can compute in code. There is
no LLM judge in the loop.
The setting is Zoomberg Brokerage, a fictional US retail brokerage. At each agent turn
the model writes the reply to the customer and a structured triage action: route or
escalate, which queue, what priority, which flags to raise, and which policies it relied on.
Every action is checked against a written 63-policy pack.
Split
Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KartiOS/fintech-support-triage.adaption-india-medical-triage-safety
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-india_medical_triage_safety
This dataset contains prompt-completion pairs for medical triage scenarios specific to India, covering emergencies like seizures, snake bites, and chest pain across various Indian languages. Each entry classifies severity, provides safe response guidance, lists unsafe actions to avoid, and specifies escalation steps such as calling emergency services. The… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-india-medical-triage-safety.medical-symptom-triage-csvhospital-triage-and-patient-history-data
Hospital Triage and Patient History Data
Tabular dataset of emergency-department triage records and patient history,
suitable for hospital-admission prediction and clinical-tabular LLM
benchmarking.
Source
This is a re-hosted copy of the dataset released by Hong, Haimovich, and
Taylor (Yale) at https://github.com/yaleemmlc/admissionprediction. The data
has been losslessly converted from the original R .RData (via an
intermediate .feather) to Apache Parquet with zstd… See the full description on the dataset page: https://huggingface.co/datasets/kondratevakate/hospital-triage-and-patient-history-data.air-track-triage
AmberTrace — Air Track Triage
ISR airspace triage: certified clear/monitor/escalate decisions over synthetic radar tracks. Features and prompts only — triage answers are obtained live from AmberTrace.
AT = gold — this dataset ships no answers
The AmberTrace verifier is the answer. Every certified-answer column
(gold / oracle / decision / triage_reason / undecidable) has been
stripped from these files: a public (features → certified decision) map
would give the… See the full description on the dataset page: https://huggingface.co/datasets/AmberTraceLabs/air-track-triage.bilingual-ticket-triage-decisionstriage-medical-dataset
Dataset release
Version: 2026-03-19-v1
Published at: 2026-03-19T15:42:16+00:00
Repo: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset
Dataset Card - POC Triage Medical
Fiche unifiee: inventaire des sources, strategie de selection, schema, gouvernance.
1) Description
Dataset bilingue FR/EN pour triage medical initial.
Le pipeline produit deux artefacts principaux:
SFT: paires instruction/reponse pour le fine-tuning supervise.
DPO: paires… See the full description on the dataset page: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset.vt-unified-triage-metadataprojet14-medical-triage-dataset
Projet 14 Medical Triage Dataset
Dataset bilingue utilise dans le cadre d'un projet academique de triage medical avec Qwen3-1.7B-Base.
Il contient des donnees pour :
le fine-tuning supervise (SFT) ;
l'alignement par preferences (DPO) ;
une evaluation de securite clinique.
Fichiers
SFT
sft_train.jsonl
sft_validation.jsonl
sft_test.jsonl
Champs :
instruction : question ou cas clinique presente au modele ;
response : reponse attendue ;
source :… See the full description on the dataset page: https://huggingface.co/datasets/PCelia/projet14-medical-triage-dataset.agent-traces-customer-support-triage
Agent Traces: customer-support-triage
Synthetic multi-agent workflow traces with LLM-enriched content for the customer-support-triage domain.
Part of the juliensimon/open-agent-traces collection — 10 datasets covering diverse domains and workflow patterns.
What is this dataset?
This dataset contains 1,483 events across 50 workflow runs, each representing a complete multi-agent execution trace. Every trace includes:
Agent reasoning — chain-of-thought for each agent step… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/agent-traces-customer-support-triage.
