recognition
table-transformer-structure-recognitionwav2vec2-large-xlsr-53-gender-recognition-librispeechhubert-large-speech-emotion-recognition-russian-dusha-finetunedtable-transformer-structure-recognition-v1.1-allmeiki.txt.recognition.v0emotion-recognition-wav2vec2-IEMOCAPpedestrian_gender_recognitionwav2vec2-lg-xlsr-en-speech-emotion-recognition
synthetic-medical-document-recognition-benchmark
Synthetic Medical Document Recognition Benchmark
This dataset contains synthetic, English-language medical records rendered as
documents for evaluating automated data extraction and de-identification
systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple
visual representations derived from that record.
Every rendered document is clearly marked as synthetic. This makes the dataset
suitable for manual testing, product demonstrations, and workflows that… See the full description on the dataset page: https://huggingface.co/datasets/morzel85/synthetic-medical-document-recognition-benchmark.Openpdf-Analysis-Recognition
Openpdf-Analysis-Recognition
The Openpdf-Analysis-Recognition dataset is curated for tasks related to image-to-text recognition, particularly for scanned document images and OCR (Optical Character Recognition) use cases. It contains over 6,900 images in a structured imagefolder format suitable for training models on document parsing, PDF image understanding, and layout/text extraction tasks.
Attribute
Value
Task
Image-to-Text
Modality
Image
Format
ImageFolder… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-Analysis-Recognition.evals-speech-recognition-cy-en
Welsh ASR Model Evaluation Transcription Dataset
This resource compiles the output transcriptions from multiple Welsh Automatic Speech Recognition (ASR) models across several test sets.
The data is structured hierarchically:
Splits delineate the individual test sets.
Configs within each split detail the performance (transcriptions) of a specific ASR model and its version on that set.
Metrics Results
model
test
task
wer
cer… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/evals-speech-recognition-cy-en.Collective-Activity-Recognition
Annotation Format
Every 10th frame in all video sequences was manually annotated with the following information for each detected person:
Bounding box location
Activity class
Pose direction
Annotation Fields
Each annotation follows the format:
<frame_number> <x> <y> <width> <height> <class_id> <pose_id>
Field
Description
frame_number
Frame identifier
x
X-coordinate of the bounding box (top-left corner)
y
Y-coordinate of the bounding box (top-left… See the full description on the dataset page: https://huggingface.co/datasets/litforth/Collective-Activity-Recognition.MJSynth_text_recognition
Dataset Card for "MJSynth_text_recognition"
This is the MJSynth dataset for text recognition on document images, synthetically generated, covering 90K English words.
It includes training, validation and test splits.
Source of the dataset: https://www.robots.ox.ac.uk/~vgg/data/text/
Use dataset streaming functionality to try out the dataset quickly without downloading the entire dataset (refer: https://huggingface.co/docs/datasets/stream)
Citation details provided on the source… See the full description on the dataset page: https://huggingface.co/datasets/priyank-m/MJSynth_text_recognition.TACO-Waste-Recognition
