Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QIRIM /crh-parallel-corpora-document-level-noisytabulartranslation10K<n<100K1 likes52 downloads2y agoHugging Face02passport-visa-photo-studio /document-photo-requirements Verified Document Photo Requirements Dataset A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact. Dataset summary Version: 1.0.0 Release date: 2026-08-14 Latest source review represented: 2026-08-11 Records: 18 (12 passport, 5 visa, 1 national ID) Coverage: 14 countries or regions Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.tabularn<1K0 likes32 downloads2mo agoHugging Face03pankaj9296 /document-accuracy-aggregation-examples Document Accuracy Aggregation Examples Two error distributions can have the same 99% field accuracy and radically different document error rates: 1% versus 50%. This educational dataset provides reproducible, synthetic examples for understanding document AI evaluation. It helps developers test metric aggregation and explain why field accuracy alone does not describe the proportion of complete documents that need correction. How we calculated these results… See the full description on the dataset page: https://huggingface.co/datasets/pankaj9296/document-accuracy-aggregation-examples.tabular1K<n<10K0 likes26 downloads2d agoHugging Face04ClarusC64 /legal-document-version-redline-final-coherence-risk-v0.1What this dataset does You receive version history redline summary final id sent or filed id approval record mismatch flags You decide coherent or incoherent Daily use wrong attachment prevention filing version QC approval gap detection tabulartext-classificationn<1K0 likes18 downloads8mo agoHugging Face05abhinavdread /msme-dispute-document-corpus MSME Dispute Document Corpus (Synthetic OCR) Dataset Description This dataset contains 8,000+ synthetic document samples designed to train AI models for the Indian MSME (Micro, Small, and Medium Enterprises) dispute resolution sector. It is specifically engineered to handle Real-World OCR Noise and Adversarial Edge Cases (e.g., distinguishing a "Proforma Invoice" from a valid "Tax Invoice"). The data mimics the messy, unstructured text often found in scanned PDFs, photos… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-dispute-document-corpus.tabulartext-classification1K<n<10K0 likes17 downloads8mo agoHugging Face06abhinavdread /msme-document-presence-dataset MSME Document Presence Detection Dataset Overview This dataset is designed for training binary classification models to detect the presence of mandatory documents in MSME arbitration cases using OCR-extracted text. The dataset supports automated document completeness validation systems. Each sample represents a structured arbitration case with document-specific OCR text fields and binary presence labels. Documents Covered The dataset includes detection labels… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-document-presence-dataset.tabulartext-classification10K<n<100K0 likes14 downloads8mo agoHugging Face07ClarusC64 /legal-chronology-event-document-issue-coherence-risk-v0.1What this dataset does You receive timeline summary document map issue links date checks gap flags conflict flags You decide coherent or incoherent Daily use chronology QC date conflict detection missing evidence detection gap finding tabulartext-classificationn<1K0 likes13 downloads8mo agoHugging Face08nsjain /single-document-tokenizedtabular100K<n<1M0 likes11 downloads6mo agoHugging Face09ClarusC64 /legal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does You receive doc description date author recipients privilege basis redaction choice context waiver flags You decide coherent or incoherent Daily use privilege log QC waiver risk detection disclosure challenge prep tabulartext-classificationn<1K0 likes10 downloads8mo agoHugging Face10Hadisawara /indonesian-tax-document-classification Indonesian Tax Document Classification Dataset Dataset Description Dataset ini berisi koleksi sintetis dokumen pajak Indonesia yang digunakan untuk klasifikasi jenis dokumen pajak. Dataset dirancang untuk mendukung penelitian NLP berbahasa Indonesia di bidang administrasi pajak dan pemerintahan daerah. Dataset ini dibuat berdasarkan pengalaman dan pengetahuan dari sistem administrasi pajak daerah (Bapenda), dengan struktur yang mencerminkan dokumen-dokumen nyata… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-tax-document-classification.tabulartext-classification10K<n<100K0 likes6 downloads3mo agoHugging Face11mtyrrell /NDC_documents_mastertabularn<1K0 likes5 downloads3y agoHugging Face12DeepDive-AI /AI-related-documentsgatedtabular10M<n<100M1 likes3 downloads2y agoHugging Face13Gato1777 /fragmentos_documentostabular1K<n<10K0 likes2 downloads2y agoHugging Face14Gato1777 /fragmentos_documentos_all-mpnet-base-v2tabular1K<n<10K0 likes2 downloads2y agoHugging Face15Gato1777 /fragmentos_documentos_61_all-mpnet-base-v2tabular1K<n<10K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.