datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
go-swe-bench-v0
go_swe_bench v0 — real Go bug fixes, verified by the Go toolchain
246 tasks from 79 real Go repositories. Each task is a bug-fix commit whose co-committed test is red on the
parent and green on the fix. No LLM anywhere in the build.
Mined on 2026-09-19 from the GuildLM Go mining pipeline by inverting the filter that had thrown the tests
away (the pipeline was built for SFT data; a benchmark needs the opposite). Every task was verified twice
with go test: green at the commit (≥ 1… See the full description on the dataset page: https://huggingface.co/datasets/guildlm/go-swe-bench-v0.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.guidellm-agentic-coding-trajectories
GuideLLM agentic coding trajectories
A sampled serving-load benchmark derived from Thoughtworks agentic-coding-trajectories, for GuideLLM and an OpenAI-compatible /v1/chat/completions endpoint. There are 630 rows representing 481 unique source sessions, across the same 8turn, 24turn, and 48turn configurations as the earlier version.
The configuration names now refer to original logical steps, not always HTTP request counts. Native tool steps expand into a tool-call request and a… See the full description on the dataset page: https://huggingface.co/datasets/zetomatoz/guidellm-agentic-coding-trajectories.guiowl-curated-corpus
GUI-Owl Curated Corpus
This dataset publishes the full curated mobile GUI-agent supervised fine-tuning corpus in a unified norm1000 mobile_use action format. Each row pairs a mobile UI screenshot with an instruction and a normalized target tool call for training GUI agents.
The published files are the curated parquet shards as produced by the source canonicalizers. No parquet shards are merged, re-sharded, or sampled during upload.
Sources
Source
Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/guiowl-curated-corpus.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/guidelines.EC-Guide
This repo is only used for dataset viewer. Please download from here.
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-b1/guidelines.task879_schema_guided_dstc8_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task879_schema_guided_dstc8_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task879_schema_guided_dstc8_classification.guitar_tabDataset of music tablature, in alphaTex (https://alphatab.net/docs/alphatex)
format, converted from Guitar Pro files (gp3, gp4, gp5, which are downloaded
from https://rutracker.org/forum/viewtopic.php?t=2888130GUIMid
Breaking the Data Barrier – Building GUI Agents Through Task Generalization
🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data
TODO List
Report and release the GUIMid with larger size and more domains (10th May expecetd)
1. Data Overview
AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks.
The performances of different domains as mid-training data are as follows:
Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K… See the full description on the dataset page: https://huggingface.co/datasets/georgeqiao12138/guidelines.Sora-Ecommerce-Guide
Sora Ecommerce Guide Dataset
This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks.
Splits
train: 9 samples
test: 2 samples
Features
instruction: System/task instruction context.
input: The prompt, question, or user query.
output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.mutopia_guitar_dataset
Mutopia Guitar Dataset
Dataset Summary
Mutopia guitar dataset consists of the soloist guitar pieces of the Mutopia Project. I encoded the MIDI files into text tokens using the excellent implementation of Dr. Tristan Beheren of the paper: MMM: Exploring Conditional Multi-Track Music Generation with the Transformer.
The dataset mainly contains guitar music from western classical composers, such as Sor, Aguado, Carcassi, and Giuliani.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/juancopi81/mutopia_guitar_dataset.GUIHazard
GUIHazard Dataset
GUIHazard: Evaluating GUI Agent Safety in Multi-Platform and Cross-Platform Workflows
GUIHazard is a cross-platform GUI-agent safety benchmark covering desktop, web, mobile, and cross-platform workflows. This Hugging Face repository contains the released benchmark data only.
For code, environment setup, and running scripts, please see the GitHub repository:
https://github.com/aifinlab/GUIHazard
Dataset Summary
GUIHazard evaluates whether GUI… See the full description on the dataset page: https://huggingface.co/datasets/AIFin-Lab/GUIHazard.gui_actor_webdataset
GUI-Actor WebDataset
A WebDataset format version of the GUI-Actor dataset for training vision-language models on GUI interaction tasks.
Usage
import webdataset as wds
# Load the dataset
dataset = wds.WebDataset("path/to/shards-*.tar")
dataset = dataset.decode("pilrgb").to_tuple("jpg", "json")
for image, metadata in dataset:
# Process image and metadata
pass
Citation
Please cite the original GUI-Actor paper if you use this dataset in your research.
epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining.
human-guided-superintelligence
Human-Guided Superintelligence: Commercial Infrastructure of the Safety Asset Class (SAC)
Dataset Overview
This dataset maps out the enterprise deployment architecture, licensing frameworks, and commercial integration primitives for the Safety Asset Class (SAC) ecosystem pioneered by Michael Aaron Russell.
It specifically codifies the mechanisms of Human-Guided Superintelligence—ensuring that recursively self-improving algorithmic stacks remain bounded by… See the full description on the dataset page: https://huggingface.co/datasets/PrimarchMI/human-guided-superintelligence.hulk_dataset_0.1This dataset is AFAIK (12 january 2024) the biggest ready to use open source dataset to finetune LLMs. It contains more than 3.8 million chat samples.
Its a collection of multiple different datasets. Some of them have been built using GPT4 or using scraped data. Here is the list:
gathnex/Gath_baize
teknium/openhermes
nomic-ai/gpt4all-j-prompt-generations
teknium/dataforge-economics
Anthropic/hh-rlhf: we kept only the selected prompts
teknium1_GPTeacher_codegen… See the full description on the dataset page: https://huggingface.co/datasets/guigux/hulk_dataset_0.1.clinical-guideline-strength-correspondence-v0.1
What this dataset tests
Guideline strength must track evidence strength.
Authority must not exceed data.
Why it exists
Guidelines often harden too early.
Language outruns certainty.
This set checks whether recommendation force matches evidence quality.
Data format
Each row contains
evidence_profile
guideline_recommendation
strength_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
evidence_profile… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-guideline-strength-correspondence-v0.1.epfl-llm_guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/minsu/epfl-llm_guidelines.paragen-security-sft-alpaca
paragen-security-sft-alpaca
Alpaca-format instruction-tuning data used to train the Vanilla security
baseline (and as the source for the tokenized multi-stream cache used to
train the Stream(Ours) security checkpoint) in the paragen_llm
security/prompt-injection-robustness experiments (Table 3: TensorTrust,
Gandalf, Purple, RuLES, StruQ, NESSiE, IFEval).
Format: JSONL, one object per line, fields instruction / input / output
(standard Alpaca schema).
Size: 48,538 examples.… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/paragen-security-sft-alpaca.japan-math-philosophy-prompts
Japan Math Philosophy Prompts
Microdataset autoral com problemas que combinam matemática e reflexão
filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em
pt-BR, en e ja e mantida integralmente no split train.
Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As
respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos
indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.tabiji-travel-safety-guides
Tabiji Travel & Safety Guides
AI-curated travel data from tabiji.ai: destination profiles, day-by-day itineraries, head-to-head comparisons, safety profiles, country-level travel advisories, and city-level scam guides — sourced from Reddit, government advisories (US State Dept., UK FCDO), and editorial curation.
What's in here
Config
Records
Description
destinations
6,498
Global destination catalog: climate, currency, language, plug type, tap-water safety… See the full description on the dataset page: https://huggingface.co/datasets/tabiji/tabiji-travel-safety-guides.gui_grounding_dataset-100
Supported Tasks
Natural Language → GUI Action Grounding
Convert user instructions into JSON action objects.
Instruction Following
Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”).
Multi-step UI Automation
Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot).
Languages
English (en)
Generated with simple variations (synonyms, phrasings).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-100.gui_grounding_dataset-1k
Supported Tasks
Natural Language → GUI Action Grounding
Convert user instructions into JSON action objects.
Instruction Following
Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”).
Multi-step UI Automation
Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot).
Languages
English (en)
Generated with simple variations (synonyms, phrasings).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-1k.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/guidelines.pulaar_corpus
Ndimaagu Pulaar Corpus
Dataset Description
This dataset contains the full transcription of the Pulaar folktale "Ndimaagu" (Nobleness/Dignity). It follows the story of Daado, Yero, and the challenges they face regarding honor and loyalty.
[cite_start]Source: ndimaagu.pdf [cite: 384, 544]
[cite_start]Language: Pulaar (ff) [cite: 384]
Format: Apache Parquet
Structure
Each entry in the dataset represents a narrative segment or a dialogue:
[cite_start]id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/guizme/pulaar_corpus.smolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
iceland-tech-christian-ethics-prompts
Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts
This microdataset contains 24 original discussion prompts arranged as 12
parallel pt-BR/English pairs. Each explicitly fictional scenario combines a
landscape motif inspired by Iceland, a technology-governance dilemma, and
concepts that may be explored through Christian ethics. The records do not
describe real Icelandic institutions, policies, communities, or practices, and
they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.
