datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
crh-parallel-corpora-document-level-noisydocument-photo-requirements
Verified Document Photo Requirements Dataset
A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact.
Dataset summary
Version: 1.0.0
Release date: 2026-08-14
Latest source review represented: 2026-08-11
Records: 18 (12 passport, 5 visa, 1 national ID)
Coverage: 14 countries or regions
Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.document-accuracy-aggregation-examples
Document Accuracy Aggregation Examples
Two error distributions can have the same 99% field accuracy and radically different document error rates: 1% versus 50%.
This educational dataset provides reproducible, synthetic examples for understanding document AI evaluation. It helps developers test metric aggregation and explain why field accuracy alone does not describe the proportion of complete documents that need correction.
How we calculated these results… See the full description on the dataset page: https://huggingface.co/datasets/pankaj9296/document-accuracy-aggregation-examples.legal-document-version-redline-final-coherence-risk-v0.1What this dataset does
You receive
version history
redline summary
final id
sent or filed id
approval record
mismatch flags
You decide
coherent
or
incoherent
Daily use
wrong attachment prevention
filing version QC
approval gap detection
msme-dispute-document-corpus
MSME Dispute Document Corpus (Synthetic OCR)
Dataset Description
This dataset contains 8,000+ synthetic document samples designed to train AI models for the Indian MSME (Micro, Small, and Medium Enterprises) dispute resolution sector.
It is specifically engineered to handle Real-World OCR Noise and Adversarial Edge Cases (e.g., distinguishing a "Proforma Invoice" from a valid "Tax Invoice"). The data mimics the messy, unstructured text often found in scanned PDFs, photos… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-dispute-document-corpus.msme-document-presence-dataset
MSME Document Presence Detection Dataset
Overview
This dataset is designed for training binary classification models to detect the presence of mandatory documents in MSME arbitration cases using OCR-extracted text.
The dataset supports automated document completeness validation systems.
Each sample represents a structured arbitration case with document-specific OCR text fields and binary presence labels.
Documents Covered
The dataset includes detection labels… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-document-presence-dataset.legal-chronology-event-document-issue-coherence-risk-v0.1What this dataset does
You receive
timeline summary
document map
issue links
date checks
gap flags
conflict flags
You decide
coherent
or
incoherent
Daily use
chronology QC
date conflict detection
missing evidence detection
gap finding
single-document-tokenizedlegal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does
You receive
doc description
date
author
recipients
privilege basis
redaction choice
context
waiver flags
You decide
coherent
or
incoherent
Daily use
privilege log QC
waiver risk detection
disclosure challenge prep
indonesian-tax-document-classification
Indonesian Tax Document Classification Dataset
Dataset Description
Dataset ini berisi koleksi sintetis dokumen pajak Indonesia yang digunakan untuk klasifikasi jenis dokumen pajak. Dataset dirancang untuk mendukung penelitian NLP berbahasa Indonesia di bidang administrasi pajak dan pemerintahan daerah.
Dataset ini dibuat berdasarkan pengalaman dan pengetahuan dari sistem administrasi pajak daerah (Bapenda), dengan struktur yang mencerminkan dokumen-dokumen nyata… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-tax-document-classification.NDC_documents_masterAI-related-documentsfragmentos_documentosfragmentos_documentos_all-mpnet-base-v2fragmentos_documentos_61_all-mpnet-base-v2
