datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Unified_Document_Understanding_Dataset
Dataset Card for Read-Parsing-Describe: Unified Scientific Document Understanding
Read-Parsing-Describe (Unified Scientific Document Understanding) is a pioneering multimodal benchmark designed to train and evaluate models on the complex structures of scientific documents. Unlike traditional document datasets that treat visual elements merely as isolated layout blocks, RPD transforms document parsing into an accessibility-driven, cross-modal grounding task.
🌟 Key… See the full description on the dataset page: https://huggingface.co/datasets/chen-doc-ai/Unified_Document_Understanding_Dataset.sroie_document_understanding
Dataset Card for "sroie_document_understanding"
Dataset Description
This dataset is an enriched version of SROIE 2019 dataset with additional labels for line descriptions and line totals for OCR and layout understanding.
Dataset Structure
DatasetDict({
train: Dataset({
features: ['image', 'ocr'],
num_rows: 652
})
})
Data Fields
{
'image': PIL Image object,
'ocr': [
# text box 1
{
'box':… See the full description on the dataset page: https://huggingface.co/datasets/arvindrajan92/sroie_document_understanding.Scaffolding-Design-And-Load-Specification-Document-Understanding-Benchmark
Scaffolding Design and Load Specification Document Understanding Benchmark
This benchmark focuses on web text covering scaffolding design guidance, load classifications, and configuration restrictions. Its questions test whether a model can interpret relationships among numerical conditions, applicable configurations, and limitations using the source text. Each record includes a source document, question, reference answer, supporting evidence, and a summary of constraint… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Scaffolding-Design-And-Load-Specification-Document-Understanding-Benchmark.
