datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tarakanov-notesmy_notes
My Notes 📓
This repository contains my lecture notes from graduate school on following topics 👇🏼
Data Science: 8 cheatsheets
Machine Learning (follows Tom Mitchell's book): 25 pages of notes
Statistics: 9 cheatsheets
Deep Learning: 12 cheatsheets, will upload more
Image Processing (follows digital image processing book): 21 cheatsheets
Data Structures and Algorithms (follows this book by Goodrich): 26 cheatsheets
✨ Some notes ✨
Most of these notes aren't intended to teach a… See the full description on the dataset page: https://huggingface.co/datasets/merve/my_notes.fast-vibevoice-notesaugmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.English-Handwritten-Math-Notes-Dataset
English Handwritten Math Notes Dataset
This dataset contains high-resolution images of handwritten mathematical notes written in English. It includes problem statements, worked examples, formulas, and annotated derivations. The dataset supports AI research in handwriting recognition, mathematical OCR, and document understanding for STEM applications.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/English-Handwritten-Math-Notes-Dataset.rednote-xiaohongshu-notes
RedNote (Xiaohongshu) Notes with Engagement and Save Rates
10,427 posts from RedNote (小红书 / Xiaohongshu), across 8 content verticals, with full Chinese post text, four separate engagement metrics, and province-level geography.
Most social datasets give you text and a like count. RedNote separates saves from likes, and that single distinction turns out to measure something the like count cannot.
The finding this dataset exists for
A post can be useful or it can be… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/rednote-xiaohongshu-notes.Handwritten-Physics-Notes-Dataset
English Handwritten Physics Notes Dataset
This dataset contains high-resolution images of handwritten physics notes written in English. The collection includes theoretical explanations, formulas, diagrams, derivations, and problem-solving steps. It is designed to support AI research in handwriting recognition, scientific OCR, and document understanding for physics and STEM education.
Contact
For queries or collaborations related to this dataset, contact:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Physics-Notes-Dataset.Handwritten-Computer-Science-Notes-Dataset
English Handwritten Computer Science Notes Dataset
This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education.
Contact
For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.synthetic_clinical_notes
Synthetic Clinical Notes
This dataset contains synthetic data from our synthetic clinical notes pipeline. You can find out more on our GitHub.
⚠️ Important Notice to Users ⚠️
All data found in this repository is entirely synthetic.
Synthetic data is artificially generated data that mimics real-world data. It is typically created using real data as a seed and adding noise. However, in our pipeline no real data is used at any point. Synthetic data can help with… See the full description on the dataset page: https://huggingface.co/datasets/NHSEDataScience/synthetic_clinical_notes.equational-theory-sair-notesEvery notes row now includes original_question_text (list<string>). The texts are copied verbatim from the original source questions/tasks; there are no null, blank, synthesized or truncated replacements. Bank lists follow related_question_ids in the same order. Initial Equational R0 notes without a question ID contain the exact assigned task recovered from their saved source-session transcript. In rollout notes tables, the list contains the exact original task for that row’s task_id. Every… See the full description on the dataset page: https://huggingface.co/datasets/U-WIN/equational-theory-sair-notes.Handwritten-Chemistry-Notes-Dataset
English Handwritten Chemistry Notes Dataset
This dataset contains high-resolution images of handwritten chemistry notes written in English. The collection includes equations, reaction mechanisms, periodic table references, structural diagrams, and descriptive explanations. It supports AI research in handwriting recognition, chemical structure understanding, and document analysis for STEM and educational applications.
Contact
For queries or collaborations related to this… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Chemistry-Notes-Dataset.x-community-notes-parquet-20250222All Twitter/X Community Notes data converted to Parquet.
https://communitynotes.x.com/guide/en/about/introduction
Pulled Feb 22, 2025
Handwritten-Biology-Notes-Dataset
English Handwritten Biology Notes Dataset
This dataset contains high-resolution images of handwritten biology notes written in English. The collection includes labeled diagrams, definitions, explanations of biological processes, and annotated sketches. It supports AI research in handwriting recognition, diagram understanding, and document interpretation within the field of life sciences.
Contact
For queries or collaborations related to this dataset, contact:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Biology-Notes-Dataset.icd10-clinical-notes
ICD-10 Multilingual Clinical Notes Dataset
A comprehensive multilingual dataset of ICD-10 diagnosis codes with clinical journal notes in 34 languages.
Author: Birger Moëll, Department of Linguistics and Philology, Uppsala University
Dataset Description
This dataset provides ICD-10 codes with:
Official diagnosis names in 34 languages (24 EU + 10 major world languages)
Sample clinical journal notes (English and Swedish)
Train/test splits for classifier training… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/icd10-clinical-notes.community-notes-br
Community Notes / X — Snapshot Público
Dataset estruturado a partir dos dumps públicos do sistema Community Notes (antigo Birdwatch) da plataforma X (antigo Twitter).
Motivação
O Community Notes é um sistema de moderação colaborativa onde usuários voluntários escrevem notas contextuais sobre publicações potencialmente enganosas e avaliam as notas de outros participantes. Um algoritmo de consenso determina quais notas são exibidas publicamente. Este dataset… See the full description on the dataset page: https://huggingface.co/datasets/histlearn/community-notes-br.ICDAR_2025_Handwritten_Notes_Understanding_Challengeso100_all_notes_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 40,
"total_frames": 11553,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Cidoyi/so100_all_notes_1.Korean_Handwritten_Notes_Dataset
Korean Handwritten Notes Dataset
This dataset contains high-resolution images of Korean handwritten notes, including personal notes, class notes, and informal writings. The dataset has been anonymized and curated to support AI research in handwriting recognition, OCR, and document understanding.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported Tasks
Task Categories:
Image… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Korean_Handwritten_Notes_Dataset.teacher-notes-severity
teacher-notes-severity
Synthetic dataset for training a local student-notes severity classifier (fine-tuned from Qwen/Qwen3-1.7B).
Each row is a chat-formatted prompt/completion pair: the user turn is a teacher note (single, compound,
or a cumulative running log), the assistant turn is a JSON label
{"category": "commendation|misbehavior|academic_concern", "severity": <int>, "escalate": <bool>}.
Severity scale: commendations are negative (-95..-10); routine notes 5-55; serious… See the full description on the dataset page: https://huggingface.co/datasets/jeremierostan/teacher-notes-severity.crowdsourced-notesguitar-fretboard-notes
Guitar Single-Note Recordings
A dataset of 390 single-note guitar recordings spanning 6 strings and frets 0-12, recorded by two players on acoustic and electric guitars.
Dataset Summary
This dataset contains isolated single-note recordings from a standard-tuned guitar. Each recording captures one note played on a specific string and fret combination, covering the first 12 frets across all 6 strings (78 unique notes per source). The recordings are raw, unprocessed 44100 Hz… See the full description on the dataset page: https://huggingface.co/datasets/collegefishiesd/guitar-fretboard-notes.SLR-NoteSense
SLR-NoteSense
SLR-NoteSense is a dual-sensor red-green-blue-clear (RGBC) dataset for Sri Lankan banknote denomination recognition. The seven-denomination update adds the LKR 2000 class to the original six-denomination dataset.
The dataset contains 14,846 paired acquisition rows from 743 distinct physical banknotes, covering LKR 20, 50, 100, 500, 1000, 2000, and 5000. After applying the documented sensor-settling rule, the stable view contains 14,104 paired measurement rows.… See the full description on the dataset page: https://huggingface.co/datasets/vennsa/SLR-NoteSense.Cedar-Field-Notes-Raw
Cedar Field Notes — Raw Release
This card preserves the field observations before normalization.
Upstream data
The records were adapted from the Cedar Survey Archive. Its authoritative release record states: Open Data Commons Attribution License.
Companion release
The derived card is Cedar Field Notes Derived.
Provenance
The source attribution is retained in the release history.
Clearance record
This release may be reused under: Open Data Commons Attribution… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Cedar-Field-Notes-Raw.Cedar-Field-Notes-Derived
Cedar Field Notes — Derived Release
This card contains transcribed observations normalized with a language model.
Upstream model
The transformation used the CedarText-7B model. Its authoritative release record states: Open Data Commons Attribution License.
Dataset
The raw companion card is Cedar Field Notes Raw.
Editorial note
Examples are retained for provenance review.
Clearance record
This release may be reused under: Open Data Commons Attribution License
omlx-m1-ultra-tuning-notes
Also on GitHub: https://github.com/6667084/omlx-m1-ultra-tuning-notes
English | 简体中文
Tuning 5 local models on an M1 Ultra (64 GB) with oMLX 0.7.0 — measured notes
Tested 2026-10-06 / 10-07 (10-07 additions: section 9 on code-model candidates, section 10 on hot cache and concurrency). Every number comes from one machine. Every adopted change was checked with
ABBA runs in the same session and a fixed 260-question quality set. Community numbers are labelled.
Update 2026-10-07… See the full description on the dataset page: https://huggingface.co/datasets/YCF-AI/omlx-m1-ultra-tuning-notes.ai-waf-dataset
Synthetic HTTP Requests Dataset for AI WAF Training
This dataset is synthetically generated and contains a diverse set of HTTP requests, labeled as either 'benign' or 'malicious'. It is designed for training and evaluating Web Application Firewalls (WAFs), particularly those based on AI/ML models.
The dataset aims to provide a comprehensive collection of both common and sophisticated attack vectors, alongside a wide array of legitimate traffic patterns.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/ai-waf-dataset.household-notes
Household and Hub reading notes
A small documentation corpus, not a benchmark.
household-note: Simplified Chinese how-tos. Each row is one failure mode (limescale vs soap scum, leftover rice, washer gasket, etc.). Diagnose first, then a short procedure, including when to stop.
hub-reading-note: English notes about using the Hugging Face Hub (model cards, safetensors, datasets, Spaces, revisions).
Use it to try load_dataset, RAG demos, or tokenizer tests. Do not treat it as… See the full description on the dataset page: https://huggingface.co/datasets/bianbian888/household-notes.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.mac-app-store-apps-release-notes
Dataset Card for Macappstore Applications Release Notes
📌 Dataset status: static snapshot (no scheduled updates). This dataset is derived from the December 2023 – January 2024 Mac App Store metadata snapshot and reflects the store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications release notes extracted from the metadata from the public API.
Curated by: MacPaw Way Ltd.
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-release-notes.
