datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fast-vibevoice-notesaugmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.rednote-xiaohongshu-notes
RedNote (Xiaohongshu) Notes with Engagement and Save Rates
10,427 posts from RedNote (小红书 / Xiaohongshu), across 8 content verticals, with full Chinese post text, four separate engagement metrics, and province-level geography.
Most social datasets give you text and a like count. RedNote separates saves from likes, and that single distinction turns out to measure something the like count cannot.
The finding this dataset exists for
A post can be useful or it can be… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/rednote-xiaohongshu-notes.equational-theory-sair-notesEvery notes row now includes original_question_text (list<string>). The texts are copied verbatim from the original source questions/tasks; there are no null, blank, synthesized or truncated replacements. Bank lists follow related_question_ids in the same order. Initial Equational R0 notes without a question ID contain the exact assigned task recovered from their saved source-session transcript. In rollout notes tables, the list contains the exact original task for that row’s task_id. Every… See the full description on the dataset page: https://huggingface.co/datasets/U-WIN/equational-theory-sair-notes.x-community-notes-parquet-20250222All Twitter/X Community Notes data converted to Parquet.
https://communitynotes.x.com/guide/en/about/introduction
Pulled Feb 22, 2025
icd10-clinical-notes
ICD-10 Multilingual Clinical Notes Dataset
A comprehensive multilingual dataset of ICD-10 diagnosis codes with clinical journal notes in 34 languages.
Author: Birger Moëll, Department of Linguistics and Philology, Uppsala University
Dataset Description
This dataset provides ICD-10 codes with:
Official diagnosis names in 34 languages (24 EU + 10 major world languages)
Sample clinical journal notes (English and Swedish)
Train/test splits for classifier training… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/icd10-clinical-notes.community-notes-br
Community Notes / X — Snapshot Público
Dataset estruturado a partir dos dumps públicos do sistema Community Notes (antigo Birdwatch) da plataforma X (antigo Twitter).
Motivação
O Community Notes é um sistema de moderação colaborativa onde usuários voluntários escrevem notas contextuais sobre publicações potencialmente enganosas e avaliam as notas de outros participantes. Um algoritmo de consenso determina quais notas são exibidas publicamente. Este dataset… See the full description on the dataset page: https://huggingface.co/datasets/histlearn/community-notes-br.teacher-notes-severity
teacher-notes-severity
Synthetic dataset for training a local student-notes severity classifier (fine-tuned from Qwen/Qwen3-1.7B).
Each row is a chat-formatted prompt/completion pair: the user turn is a teacher note (single, compound,
or a cumulative running log), the assistant turn is a JSON label
{"category": "commendation|misbehavior|academic_concern", "severity": <int>, "escalate": <bool>}.
Severity scale: commendations are negative (-95..-10); routine notes 5-55; serious… See the full description on the dataset page: https://huggingface.co/datasets/jeremierostan/teacher-notes-severity.guitar-fretboard-notes
Guitar Single-Note Recordings
A dataset of 390 single-note guitar recordings spanning 6 strings and frets 0-12, recorded by two players on acoustic and electric guitars.
Dataset Summary
This dataset contains isolated single-note recordings from a standard-tuned guitar. Each recording captures one note played on a specific string and fret combination, covering the first 12 frets across all 6 strings (78 unique notes per source). The recordings are raw, unprocessed 44100 Hz… See the full description on the dataset page: https://huggingface.co/datasets/collegefishiesd/guitar-fretboard-notes.ai-waf-dataset
Synthetic HTTP Requests Dataset for AI WAF Training
This dataset is synthetically generated and contains a diverse set of HTTP requests, labeled as either 'benign' or 'malicious'. It is designed for training and evaluating Web Application Firewalls (WAFs), particularly those based on AI/ML models.
The dataset aims to provide a comprehensive collection of both common and sophisticated attack vectors, alongside a wide array of legitimate traffic patterns.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/ai-waf-dataset.household-notes
Household and Hub reading notes
A small documentation corpus, not a benchmark.
household-note: Simplified Chinese how-tos. Each row is one failure mode (limescale vs soap scum, leftover rice, washer gasket, etc.). Diagnose first, then a short procedure, including when to stop.
hub-reading-note: English notes about using the Hugging Face Hub (model cards, safetensors, datasets, Spaces, revisions).
Use it to try load_dataset, RAG demos, or tokenizer tests. Do not treat it as… See the full description on the dataset page: https://huggingface.co/datasets/bianbian888/household-notes.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.Alcohol_Use_Clinical_Notes_GPT4Contributions: The dataset was created by Dr. Uri Kartoun.
Use Case: Leveraging Large Language Models for Enhanced Clinical Narrative Analysis: An Application in Alcohol Use Detection
Dataset Summary: This dataset contains 1,500 samples of expressions indicating alcohol use or its negation, generated from clinical narrative notes using OpenAI's ChatGPT 4 model. It's designed to support NLP applications that require the identification of alcohol use references in healthcare records.
Text… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Alcohol_Use_Clinical_Notes_GPT4.mac-app-store-apps-release-notes
Dataset Card for Macappstore Applications Release Notes
📌 Dataset status: static snapshot (no scheduled updates). This dataset is derived from the December 2023 – January 2024 Mac App Store metadata snapshot and reflects the store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications release notes extracted from the metadata from the public API.
Curated by: MacPaw Way Ltd.
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-release-notes.augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.mimic-iv-notesaugmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.teacher-notes-severity-v2
Teacher Notes Severity v2
Synthetic teacher notes labeled with category, location (in_class / outside_class), severity (-100..100, >=60 escalates) and escalate flag. 3800 train / 400 test. Generator: generate_dataset_v2.py in this repo. Continues the v1 dataset (jeremierostan/teacher-notes-severity); schema adds location.
synthetic-clinical-notes-embedded
Synthetic Clinical Notes
This dataset is post-processed version of starmpcc/Asclepius-Synthetic-Clinical-Notes:
Turn into Alpaca format (instruction, input, and output)
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
158k
Token Count
648m
Origin
https://figshare.com/authors/Zhengyun_Zhao/16480335
Source of raw data
PubMed Central (PMC) and MIMIC 3
Processing details
original, paper
Embedding Model… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/synthetic-clinical-notes-embedded.dyspnea-clinical-notes
Dataset description
This dataset contains the train unannotated clinical notes for the CRF:filling Shared Task at CL4Health2026.
The clinical notes have been collected, anonymized and annotated at the San Giovanni Bosco (SGB) hospital, Turin, Italy.
There are two splits, each representing a different language: en (English) and it (Italian). English data has been automatically translated from Italian.
Each example (2667 in total) in the dataset is composed by:
document_id: clinical… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/dyspnea-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Fadil369/augmented-clinical-notes.medical_notes_30k_questionsWe asked Minimax/Minimax-M3 to create 100 questions for the dataset https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.
The questions are created to cover multiple notes and trace their provenance through evidences and note IDs.
field-notes-transcription-packet
Field-notes Transcription Packet
Verified field-note transcriptions retained for the research archive.
Retained notes: 6
Featured note: NOTE-1002 — Cedar / English
Transcription window: 2023-01-18T16:20:00Z to 2023-12-15T10:00:00Z
Site counts (Cedar/Lark/Morrow/Pine): 2/2/1/1
Mean word counts (Cedar/Lark/Morrow/Pine): 422.5/362.5/360.0/480.0
R2E-Gym-Javi-40pct-4274-NoTestaugmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/minidiablo05/augmented-clinical-notes.icd10-clinical-notes
ICD-10 Multilingual Clinical Notes Dataset
A comprehensive multilingual dataset of ICD-10 diagnosis codes with clinical journal notes in 34 languages.
Author: Birger Moëll, Department of Linguistics and Philology, Uppsala University
Dataset Description
This dataset provides ICD-10 codes with:
Official diagnosis names in 34 languages (24 EU + 10 major world languages)
Sample clinical journal notes (English and Swedish)
Train/test splits for classifier training… See the full description on the dataset page: https://huggingface.co/datasets/Danishaqil/icd10-clinical-notes.JEP-EU-AI-Act-Mapping-Notes
JEP — EU AI Act Mapping Notes
Exploratory / non-normative / not legal advice
This dataset records limited research notes about where JEP-style signed event
structures may be relevant to documentation, traceability, or review workflows
discussed in the EU AI Act.
It does not claim that JEP provides compliance, satisfies any legal
requirement, or replaces legal, regulatory, technical, or organizational
controls.
Protocol source
JEP-Core 0.7 / draft-07:… See the full description on the dataset page: https://huggingface.co/datasets/hjs-spec/JEP-EU-AI-Act-Mapping-Notes.dutch_nursing_home_notes
dutch_nursing_home_notes
Description
This dataset was previously named dutch_nursing_home_records. The name was changed for consistency, as the term ‘notes’ is more commonly used in the community of clinical NLP
This dataset contains synthetic healthcare data generated for NLP experiments.
It mimics real-world client notes of nursing care homes for machine learning and data analysis.
Data Generation
The script describing the data generation can be found here:… See the full description on the dataset page: https://huggingface.co/datasets/ekrombouts/dutch_nursing_home_notes.
