datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zamai-pashto-documents
ZamAI-Pashto Documents
This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language.
Project Structure
data/: Contains scanned documents, extracted text, translations, and summaries.
annotations/: OCR bounding boxes, handwriting labels, and domain tags.
scripts/: OCR processing, text cleaning, and translation alignment scripts.
configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
vietnam-legal-documentseasylaw_kr_documentscrh-parallel-corpora-document-level-noisyUN_Documents_2000_2023plazos-conservacion-documentos-laborales-espana-RRHH
Plazos de conservación de documentos laborales en España (RRHH)
Tabla estructurada con los plazos de conservación que fija la normativa española para los documentos de personal: registro de jornada, nóminas, documentación de Seguridad Social, contabilidad, documentación tributaria, canal de denuncias, prevención de riesgos y datos personales. Cada fila indica el plazo, qué tipo de plazo es (mínimo, máximo, supresión obligatoria o sin plazo expreso), desde cuándo cuenta, la norma… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/plazos-conservacion-documentos-laborales-espana-RRHH.zamai-pashto-documents
Pashto
Languages: psLicense: cc-by-4.0Task categories: visual-document-retrievalSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for visual-document-retrieval tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-documents")
print(dataset)
Citation
@misc{zamai_pashto_data,
title = {{Pashto}},
author = {ZamAI / Yaqoob… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-documents.document-photo-requirements
Verified Document Photo Requirements Dataset
A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact.
Dataset summary
Version: 1.0.0
Release date: 2026-08-14
Latest source review represented: 2026-08-11
Records: 18 (12 passport, 5 visa, 1 national ID)
Coverage: 14 countries or regions
Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.transformers_documentation_endocument-accuracy-aggregation-examples
Document Accuracy Aggregation Examples
Two error distributions can have the same 99% field accuracy and radically different document error rates: 1% versus 50%.
This educational dataset provides reproducible, synthetic examples for understanding document AI evaluation. It helps developers test metric aggregation and explain why field accuracy alone does not describe the proportion of complete documents that need correction.
How we calculated these results… See the full description on the dataset page: https://huggingface.co/datasets/pankaj9296/document-accuracy-aggregation-examples.Legal-Documentwatsonx-docs-document-type
Watsonx Docs Document Type Classification
This dataset is a balanced binary document-level classification subset derived
from ibm-research/watsonxDocsQA.
Task
Classify IBM Watsonx documentation pages by their dominant user-facing purpose:
conceptual: documents primarily used to understand or look up information.
how-to: documents primarily used to complete a procedure or fix a problem.
Splits
Split
conceptual
how-to
Total
train
140
140
280… See the full description on the dataset page: https://huggingface.co/datasets/itsjhuang/watsonx-docs-document-type.legal-document-version-redline-final-coherence-risk-v0.1What this dataset does
You receive
version history
redline summary
final id
sent or filed id
approval record
mismatch flags
You decide
coherent
or
incoherent
Daily use
wrong attachment prevention
filing version QC
approval gap detection
maritime_documents_tags_classificationData Understanding and Preparation
Data Collection
BIMCO Contracts and Clauses
URL: BIMCO Contracts and Clauses
Description: This website provides a wide range of standardized contracts and clauses commonly used in the maritime industry. The documents were downloaded and used as part of our dataset, offering detailed insights into industry-specific terminology and structured data.
Data Description
Documents Collected: A total of 217 documents were collected… See the full description on the dataset page: https://huggingface.co/datasets/pranay27sy/maritime_documents_tags_classification.msme-dispute-document-corpus
MSME Dispute Document Corpus (Synthetic OCR)
Dataset Description
This dataset contains 8,000+ synthetic document samples designed to train AI models for the Indian MSME (Micro, Small, and Medium Enterprises) dispute resolution sector.
It is specifically engineered to handle Real-World OCR Noise and Adversarial Edge Cases (e.g., distinguishing a "Proforma Invoice" from a valid "Tax Invoice"). The data mimics the messy, unstructured text often found in scanned PDFs, photos… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-dispute-document-corpus.msme-document-presence-dataset
MSME Document Presence Detection Dataset
Overview
This dataset is designed for training binary classification models to detect the presence of mandatory documents in MSME arbitration cases using OCR-extracted text.
The dataset supports automated document completeness validation systems.
Each sample represents a structured arbitration case with document-specific OCR text fields and binary presence labels.
Documents Covered
The dataset includes detection labels… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-document-presence-dataset.DocumentCreationmedical-documentation-datasetlegal-chronology-event-document-issue-coherence-risk-v0.1What this dataset does
You receive
timeline summary
document map
issue links
date checks
gap flags
conflict flags
You decide
coherent
or
incoherent
Daily use
chronology QC
date conflict detection
missing evidence detection
gap finding
SoICT-Hackathon-2024-Legal-Document-RetrievalMNLP_M3_rag_documentstactics-documentsMNLP_M2_rag_documentsMNLP_M3_rag_documentFinancial_Documents_datasetadwaittagalpallewar_medical-document-ocr-text-dataset
Medical Document OCR Text Dataset
Synthetic OCR-extracted text from medical documents for NLP classification
Dataset Info
Source: Kaggle
Original Size: 22.39 MB
Kaggle Downloads: 33
Files: 1
Files
medical_documents_dataset.csv
Mirrored from Kaggle
single-document-tokenizedRETRIEVED_DOCUMENTSlegal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does
You receive
doc description
date
author
recipients
privilege basis
redaction choice
context
waiver flags
You decide
coherent
or
incoherent
Daily use
privilege log QC
waiver risk detection
disclosure challenge prep
