datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Openpdf-Analysis-Recognition
Openpdf-Analysis-Recognition
The Openpdf-Analysis-Recognition dataset is curated for tasks related to image-to-text recognition, particularly for scanned document images and OCR (Optical Character Recognition) use cases. It contains over 6,900 images in a structured imagefolder format suitable for training models on document parsing, PDF image understanding, and layout/text extraction tasks.
Attribute
Value
Task
Image-to-Text
Modality
Image
Format
ImageFolder… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-Analysis-Recognition.evals-speech-recognition-cy-en
Welsh ASR Model Evaluation Transcription Dataset
This resource compiles the output transcriptions from multiple Welsh Automatic Speech Recognition (ASR) models across several test sets.
The data is structured hierarchically:
Splits delineate the individual test sets.
Configs within each split detail the performance (transcriptions) of a specific ASR model and its version on that set.
Metrics Results
model
test
task
wer
cer… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/evals-speech-recognition-cy-en.MJSynth_text_recognition
Dataset Card for "MJSynth_text_recognition"
This is the MJSynth dataset for text recognition on document images, synthetically generated, covering 90K English words.
It includes training, validation and test splits.
Source of the dataset: https://www.robots.ox.ac.uk/~vgg/data/text/
Use dataset streaming functionality to try out the dataset quickly without downloading the entire dataset (refer: https://huggingface.co/docs/datasets/stream)
Citation details provided on the source… See the full description on the dataset page: https://huggingface.co/datasets/priyank-m/MJSynth_text_recognition.EvArEST-dataset-for-Arabic-scene-text-recognition
EvArEST
Everyday Arabic-English Scene Text dataset, from the paper: Arabic Scene Text Recognition in the Deep Learning Era: Analysis on A Novel Dataset
The dataset includes both the recognition dataset and the synthetic one in a single train and test split.
Recognition Dataset
The text recognition dataset comprises of 7232 cropped word images of both Arabic and English languages. The groundtruth for the recognition dataset is provided by a text file with each line… See the full description on the dataset page: https://huggingface.co/datasets/Melaraby/EvArEST-dataset-for-Arabic-scene-text-recognition.synth-text-recognition
Dataset Card for "Synth-Text Recognition"
This is the dataset for text recognition on document images, synthetically generated, covering 90K English words.
It includes training, validation and test splits.
text_recognition_en_zh_clean
Dataset Card for "text_recognition_en_zh_clean"
More Information needed
Military-Aircraft-Recognition-datasetThis is a remote sensing image Military Aircraft Recognition dataset that include 3842 images, 20 types, and 22341 instances annotated with horizontal bounding boxes and oriented bounding boxes.
trdg_random_en_zh_text_recognition
Dataset Card for "trdg_random_en_zh_text_recognition"
This synthetic dataset was generated using the TextRecognitionDataGenerator(TRDG) open source repo:
https://github.com/Belval/TextRecognitionDataGenerator
It contains images of text with random characters from Engilsh(en) and Chinese(zh) languages.
Reference to the documentation provided by the TRDG repo:
https://textrecognitiondatagenerator.readthedocs.io/en/latest/index.html
chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.BACHI_Chord_Recognition
BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
Paper | Project Page | Code | POP909-CL Dataset
This repository contains trained model weights and classical datasets for the paper:
Mingyang Yao, Ke Chen, Shlomo Dubnov and Taylor Berg-Kirkpatrick
"BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music."ICASSP 2026, 2025
Abstract
Automatic chord… See the full description on the dataset page: https://huggingface.co/datasets/Itsuki-music/BACHI_Chord_Recognition.Diverse-hand-gesture-recognitionnamed_entity_recognition_document_contextSROIE_2019_text_recognitionThis dataset we prepared using the Scanned receipts OCR and information extraction(SROIE) dataset.
The SROIE dataset contains 973 scanned receipts in English language.
Cropping the bounding boxes from each of the receipts to generate this text-recognition dataset resulted in 33626 images for train set and 18704 images for the test set.
The text annotations for all the images inside a split are stored in a metadata.jsonl file.
usage:
from dataset import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/priyank-m/SROIE_2019_text_recognition.taiwan-license-plate-recognitionmulti-label-food-recognition
Multi-Label Food Recognition Dataset
This is a multi-label food recognition dataset generated from single-class food images.
Each image contains 2-5 different food items composited together using natural composition methods.
Dataset Details
Total Images: 13,000
Training Images: 10,400 (80%)
Validation Images: 2,600 (20%)
Number of Classes: 90
Labels per Image: 2-5 labels
Image Format: RGB, 512x512 pixels
File Format: Parquet
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/multi-label-food-recognition.text_recognition_en_zh
Dataset Card for "text_recognition_en_zh"
More Information needed
surgical-tool-recognition-full-multiview
Surgical Tool Recognition Full Multiview
Summary
This dataset contains images of individual surgical instruments for object detection.It was originally created in YOLO format and exported here to a Hugging Face-friendly structure with metadata.jsonl files for each split.
Splits
train: 2016
validation: 252
test: 252
Total: 2520 images
Classes
0 = clamp
1 = needle_holder
2 = scalpel
3 = shear
4 = tweezer
File naming convention
Each image… See the full description on the dataset page: https://huggingface.co/datasets/joonhaim/surgical-tool-recognition-full-multiview.text-recognition
Dataset Card for "text-recognition"
More Information needed
Logo-Recognition-ResNet50-TripletNet-Embeddings-Datasetface-recognitionssi-speech-emotion-recognition
Dataset Card for SSI: Speech Emotion Recognition - Stapes AI
Dataset Details
Dataset Format for Audio Files
This is the format for the audio files in the dataset. We'll open-source the dataset soon.
Gender
M - Male
F - Female
Age Group
CH - Child (0-12)
TE - Teenager (13-19)
AD - Adult (20-60)
SE - Senior (60+)
UNK - Unknown
Utterance Type
SEN: Sentence
WOR: Word
PHR: Phrase
Sentence
DFA: "Don't Forget A… See the full description on the dataset page: https://huggingface.co/datasets/stapesai/ssi-speech-emotion-recognition.synth-text-recognition-multilines-cs
Czech Synthetic Multiline Text Recognition Dataset
A large-scale synthetic dataset for Czech multiline text recognition, containing 100,000 text images with corresponding transcriptions. Created using SynthTiger.
Dataset Description
This dataset consists of synthetically generated images of Czech text with multiple lines per image, designed for training optical character recognition (OCR) models that can handle complex multiline text layouts. Each image contains 3 lines… See the full description on the dataset page: https://huggingface.co/datasets/Empatixx/synth-text-recognition-multilines-cs.Human_Activity_RecognitionHuman Activity Recognition (HAR) using smartphones dataset. Classifying the type of movement amongst five categories:
WALKING,
WALKING_UPSTAIRS,
WALKING_DOWNSTAIRS,
SITTING,
STANDING
The experiments have been carried out with a group of 16 volunteers within an age bracket of 19-26 years. Each person performed five activities (WALKING, WALKING_UPSTAIRS, WALKING_DOWNSTAIRS, SITTING, STANDING) wearing a smartphone (Samsung Galaxy S8) in the pucket. Using its embedded accelerometer and gyroscope… See the full description on the dataset page: https://huggingface.co/datasets/DiFronzo/Human_Activity_Recognition.Quran_speech_recognition_kaggleThis dataset can be found in Kaggle
synth-text-recognition-cs
Czech Synthetic Text Recognition Dataset
A large-scale synthetic dataset for Czech text recognition, containing 454,820 text images with corresponding transcriptions. Created using SynthTiger.
Dataset Description
This dataset consists of synthetically generated images of Czech text, designed for training optical character recognition (OCR) models. Each image contains a single word or short phrase rendered with various visual effects to simulate real-world text appearance.… See the full description on the dataset page: https://huggingface.co/datasets/Empatixx/synth-text-recognition-cs.slovenian-speech-recognition
Slovenian Speech Dataset
Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.raw-food-recognition
Merged Raw Food Recognition Dataset
Dataset Description
This dataset is a comprehensive compilation of three publicly available food recognition datasets, merged and curated for raw food recognition tasks. The dataset contains images of various raw food items including fruits, vegetables, dairy products, and beverages, intended for educational purposes and the development of image recognition models.
Purpose
This dataset is created for educational purposes only… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/raw-food-recognition.american-speech-recognition-dataset
American Speech Dataset for recognition task
Dataset comprises 1,136 hours of telephone dialogues in American, collected from 1,416 native speakers across various topics and domains, achieving an impressive 95% Sentence Accuracy Rate. It is designed for research in automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/american-speech-recognition-dataset.entity_recognition
Data for the NERMuD shared task (Evalita 2023)
This data is the one used for the NERMuD shared task organized
at Evalita 2023.
The dataset contains the Wikinews, fiction, and De Gasperi subsets of KIND, where test data is used for development.
Content of the dataset
Split
Sentences
wn_train
10,912
wn_dev
2,594
wn_test
2,088
fic_train11,423
fic_dev
1,051
fic_test
1,517
adg_train
5,147
adg_dev
1,122
adg_test
521
Set
Sentences… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/entity_recognition.
