Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01prithivMLmods /Openpdf-Analysis-Recognition Openpdf-Analysis-Recognition The Openpdf-Analysis-Recognition dataset is curated for tasks related to image-to-text recognition, particularly for scanned document images and OCR (Optical Character Recognition) use cases. It contains over 6,900 images in a structured imagefolder format suitable for training models on document parsing, PDF image understanding, and layout/text extraction tasks. Attribute Value Task Image-to-Text Modality Image Format ImageFolder… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-Analysis-Recognition.imageimage-to-text1K<n<10K4 likes1.9k downloads1y agoHugging Face02techiaith /evals-speech-recognition-cy-en Welsh ASR Model Evaluation Transcription Dataset This resource compiles the output transcriptions from multiple Welsh Automatic Speech Recognition (ASR) models across several test sets. The data is structured hierarchically: Splits delineate the individual test sets. Configs within each split detail the performance (transcriptions) of a specific ASR model and its version on that set. Metrics Results model test task wer cer… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/evals-speech-recognition-cy-en.tabular100K<n<1M0 likes1.4k downloads21d agoHugging Face03priyank-m /MJSynth_text_recognition Dataset Card for "MJSynth_text_recognition" This is the MJSynth dataset for text recognition on document images, synthetically generated, covering 90K English words. It includes training, validation and test splits. Source of the dataset: https://www.robots.ox.ac.uk/~vgg/data/text/ Use dataset streaming functionality to try out the dataset quickly without downloading the entire dataset (refer: https://huggingface.co/docs/datasets/stream) Citation details provided on the source… See the full description on the dataset page: https://huggingface.co/datasets/priyank-m/MJSynth_text_recognition.imageimage-to-text1M<n<10M9 likes812 downloads3y agoHugging Face04Melaraby /EvArEST-dataset-for-Arabic-scene-text-recognition EvArEST Everyday Arabic-English Scene Text dataset, from the paper: Arabic Scene Text Recognition in the Deep Learning Era: Analysis on A Novel Dataset The dataset includes both the recognition dataset and the synthetic one in a single train and test split. Recognition Dataset The text recognition dataset comprises of 7232 cropped word images of both Arabic and English languages. The groundtruth for the recognition dataset is provided by a text file with each line… See the full description on the dataset page: https://huggingface.co/datasets/Melaraby/EvArEST-dataset-for-Arabic-scene-text-recognition.image100K<n<1M2 likes545 downloads11mo agoHugging Face05redactable-llm /synth-text-recognition Dataset Card for "Synth-Text Recognition" This is the dataset for text recognition on document images, synthetically generated, covering 90K English words. It includes training, validation and test splits. imageimage-to-text1M<n<10M4 likes529 downloads3y agoHugging Face06priyank-m /text_recognition_en_zh_clean Dataset Card for "text_recognition_en_zh_clean" More Information needed image1M<n<10M6 likes513 downloads4y agoHugging Face07Alex5666 /Military-Aircraft-Recognition-datasetThis is a remote sensing image Military Aircraft Recognition dataset that include 3842 images, 20 types, and 22341 instances annotated with horizontal bounding boxes and oriented bounding boxes. imageimage-classification1K<n<10K8 likes421 downloads3y agoHugging Face08priyank-m /trdg_random_en_zh_text_recognition Dataset Card for "trdg_random_en_zh_text_recognition" This synthetic dataset was generated using the TextRecognitionDataGenerator(TRDG) open source repo: https://github.com/Belval/TextRecognitionDataGenerator It contains images of text with random characters from Engilsh(en) and Chinese(zh) languages. Reference to the documentation provided by the TRDG repo: https://textrecognitiondatagenerator.readthedocs.io/en/latest/index.html imageimage-to-text100K<n<1M3 likes411 downloads2y agoHugging Face09priyank-m /chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition imageimage-to-text100K<n<1M36 likes360 downloads4y agoHugging Face10DaftP /Home-Assistant-requests-for-intent-detection-and-function-recognition Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.textquestion-answering100K<n<1M1 likes360 downloads6mo agoHugging Face11Itsuki-music /BACHI_Chord_Recognition BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music Paper | Project Page | Code | POP909-CL Dataset This repository contains trained model weights and classical datasets for the paper: Mingyang Yao, Ke Chen, Shlomo Dubnov and Taylor Berg-Kirkpatrick "BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music."ICASSP 2026, 2025 Abstract Automatic chord… See the full description on the dataset page: https://huggingface.co/datasets/Itsuki-music/BACHI_Chord_Recognition.textothern<1K7 likes356 downloads3mo agoHugging Face12shenasa /Diverse-hand-gesture-recognitionimage1K<n<10K0 likes349 downloads1y agoHugging Face13LegionIntel /named_entity_recognition_document_contexttabular1M<n<10M9 likes335 downloads2y agoHugging Face14priyank-m /SROIE_2019_text_recognitionThis dataset we prepared using the Scanned receipts OCR and information extraction(SROIE) dataset. The SROIE dataset contains 973 scanned receipts in English language. Cropping the bounding boxes from each of the receipts to generate this text-recognition dataset resulted in 33626 images for train set and 18704 images for the test set. The text annotations for all the images inside a split are stored in a metadata.jsonl file. usage: from dataset import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/priyank-m/SROIE_2019_text_recognition.imageimage-to-text10K<n<100K14 likes331 downloads4y agoHugging Face15EZCon /taiwan-license-plate-recognitionimageimage-segmentation1K<n<10K2 likes325 downloads1y agoHugging Face16ibrahimdaud /multi-label-food-recognition Multi-Label Food Recognition Dataset This is a multi-label food recognition dataset generated from single-class food images. Each image contains 2-5 different food items composited together using natural composition methods. Dataset Details Total Images: 13,000 Training Images: 10,400 (80%) Validation Images: 2,600 (20%) Number of Classes: 90 Labels per Image: 2-5 labels Image Format: RGB, 512x512 pixels File Format: Parquet Dataset Structure Each sample… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/multi-label-food-recognition.imageimage-classification10K<n<100K1 likes303 downloads10mo agoHugging Face17priyank-m /text_recognition_en_zh Dataset Card for "text_recognition_en_zh" More Information needed image1M<n<10M1 likes293 downloads4y agoHugging Face18joonhaim /surgical-tool-recognition-full-multiview Surgical Tool Recognition Full Multiview Summary This dataset contains images of individual surgical instruments for object detection.It was originally created in YOLO format and exported here to a Hugging Face-friendly structure with metadata.jsonl files for each split. Splits train: 2016 validation: 252 test: 252 Total: 2520 images Classes 0 = clamp 1 = needle_holder 2 = scalpel 3 = shear 4 = tweezer File naming convention Each image… See the full description on the dataset page: https://huggingface.co/datasets/joonhaim/surgical-tool-recognition-full-multiview.imageobject-detection1K<n<10K1 likes292 downloads6mo agoHugging Face19longhoang06 /text-recognition Dataset Card for "text-recognition" More Information needed image100K<n<1M0 likes290 downloads3y agoHugging Face20MahmoodAnaam /Logo-Recognition-ResNet50-TripletNet-Embeddings-Datasetimage1M<n<10M0 likes282 downloads7mo agoHugging Face21wuji3 /face-recognitionimage1M<n<10M5 likes281 downloads2y agoHugging Face22stapesai /ssi-speech-emotion-recognition Dataset Card for SSI: Speech Emotion Recognition - Stapes AI Dataset Details Dataset Format for Audio Files This is the format for the audio files in the dataset. We'll open-source the dataset soon. Gender M - Male F - Female Age Group CH - Child (0-12) TE - Teenager (13-19) AD - Adult (20-60) SE - Senior (60+) UNK - Unknown Utterance Type SEN: Sentence WOR: Word PHR: Phrase Sentence DFA: "Don't Forget A… See the full description on the dataset page: https://huggingface.co/datasets/stapesai/ssi-speech-emotion-recognition.audio10K<n<100K16 likes211 downloads2y agoHugging Face23Empatixx /synth-text-recognition-multilines-cs Czech Synthetic Multiline Text Recognition Dataset A large-scale synthetic dataset for Czech multiline text recognition, containing 100,000 text images with corresponding transcriptions. Created using SynthTiger. Dataset Description This dataset consists of synthetically generated images of Czech text with multiple lines per image, designed for training optical character recognition (OCR) models that can handle complex multiline text layouts. Each image contains 3 lines… See the full description on the dataset page: https://huggingface.co/datasets/Empatixx/synth-text-recognition-multilines-cs.image100K<n<1M0 likes193 downloads1y agoHugging Face24DiFronzo /Human_Activity_RecognitionHuman Activity Recognition (HAR) using smartphones dataset. Classifying the type of movement amongst five categories: WALKING, WALKING_UPSTAIRS, WALKING_DOWNSTAIRS, SITTING, STANDING The experiments have been carried out with a group of 16 volunteers within an age bracket of 19-26 years. Each person performed five activities (WALKING, WALKING_UPSTAIRS, WALKING_DOWNSTAIRS, SITTING, STANDING) wearing a smartphone (Samsung Galaxy S8) in the pucket. Using its embedded accelerometer and gyroscope… See the full description on the dataset page: https://huggingface.co/datasets/DiFronzo/Human_Activity_Recognition.text100K<n<1M2 likes190 downloads5y agoHugging Face25Nuwaisir /Quran_speech_recognition_kaggleThis dataset can be found in Kaggle text10K<n<100K7 likes185 downloads5y agoHugging Face26Empatixx /synth-text-recognition-cs Czech Synthetic Text Recognition Dataset A large-scale synthetic dataset for Czech text recognition, containing 454,820 text images with corresponding transcriptions. Created using SynthTiger. Dataset Description This dataset consists of synthetically generated images of Czech text, designed for training optical character recognition (OCR) models. Each image contains a single word or short phrase rendered with various visual effects to simulate real-world text appearance.… See the full description on the dataset page: https://huggingface.co/datasets/Empatixx/synth-text-recognition-cs.image100K<n<1M0 likes176 downloads1y agoHugging Face27UniDataPro /slovenian-speech-recognition Slovenian Speech Dataset Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/slovenian-speech-recognition.audioautomatic-speech-recognitionn<1K11 likes172 downloads2mo agoHugging Face28ibrahimdaud /raw-food-recognition Merged Raw Food Recognition Dataset Dataset Description This dataset is a comprehensive compilation of three publicly available food recognition datasets, merged and curated for raw food recognition tasks. The dataset contains images of various raw food items including fruits, vegetables, dairy products, and beverages, intended for educational purposes and the development of image recognition models. Purpose This dataset is created for educational purposes only… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimdaud/raw-food-recognition.imageimage-classification10K<n<100K0 likes170 downloads10mo agoHugging Face29UniDataPro /american-speech-recognition-dataset American Speech Dataset for recognition task Dataset comprises 1,136 hours of telephone dialogues in American, collected from 1,416 native speakers across various topics and domains, achieving an impressive 95% Sentence Accuracy Rate. It is designed for research in automatic speech recognition (ASR) systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in natural language processing (NLP), speech recognition, and machine… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/american-speech-recognition-dataset.textn<1K8 likes164 downloads2mo agoHugging Face30evalitahf /entity_recognition Data for the NERMuD shared task (Evalita 2023) This data is the one used for the NERMuD shared task organized at Evalita 2023. The dataset contains the Wikinews, fiction, and De Gasperi subsets of KIND, where test data is used for development. Content of the dataset Split Sentences wn_train 10,912 wn_dev 2,594 wn_test 2,088 fic_train11,423 fic_dev 1,051 fic_test 1,517 adg_train 5,147 adg_dev 1,122 adg_test 521 Set Sentences… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/entity_recognition.texttext-classification10K<n<100K0 likes162 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.