Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZamAI-Pashto /zamai-pashto-documents ZamAI-Pashto Documents This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language. Project Structure data/: Contains scanned documents, extracted text, translations, and summaries. annotations/: OCR bounding boxes, handwriting labels, and domain tags. scripts/: OCR processing, text cleaning, and translation alignment scripts. configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes663 downloads2mo agoHugging Face02ClarusC64 /maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for Triage trade doc packs before they trigger holds. You use it to flag HS code inconsistencies across documents missing certificates shipper or consignee mismatch clearance status lag not supported by doc quality Why it matters Most port delay disputes begin in paperwork. texttext-classificationn<1K1 likes100 downloads8mo agoHugging Face03YuITC /vietnam-legal-documentstext100K<n<1M1 likes54 downloads6mo agoHugging Face04jiwoochris /easylaw_kr_documentstext1K<n<10K2 likes53 downloads3y agoHugging Face05QIRIM /crh-parallel-corpora-document-level-noisytabulartranslation10K<n<100K1 likes52 downloads2y agoHugging Face06dabaisuv /UN_Documents_2000_2023text1M<n<10M1 likes35 downloads3y agoHugging Face07Nucleo360 /plazos-conservacion-documentos-laborales-espana-RRHH Plazos de conservación de documentos laborales en España (RRHH) Tabla estructurada con los plazos de conservación que fija la normativa española para los documentos de personal: registro de jornada, nóminas, documentación de Seguridad Social, contabilidad, documentación tributaria, canal de denuncias, prevención de riesgos y datos personales. Cada fila indica el plazo, qué tipo de plazo es (mínimo, máximo, supresión obligatoria o sin plazo expreso), desde cuándo cuenta, la norma… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/plazos-conservacion-documentos-laborales-espana-RRHH.texttable-question-answeringn<1K0 likes34 downloads23h agoHugging Face08tasal9 /zamai-pashto-documents Pashto Languages: psLicense: cc-by-4.0Task categories: visual-document-retrievalSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for visual-document-retrieval tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/zamai-pashto-documents") print(dataset) Citation @misc{zamai_pashto_data, title = {{Pashto}}, author = {ZamAI / Yaqoob… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes32 downloads2mo agoHugging Face09passport-visa-photo-studio /document-photo-requirements Verified Document Photo Requirements Dataset A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact. Dataset summary Version: 1.0.0 Release date: 2026-08-14 Latest source review represented: 2026-08-11 Records: 18 (12 passport, 5 visa, 1 national ID) Coverage: 14 countries or regions Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.tabularn<1K0 likes32 downloads2mo agoHugging Face10m-ric /transformers_documentation_entextn<1K0 likes27 downloads3y agoHugging Face11pankaj9296 /document-accuracy-aggregation-examples Document Accuracy Aggregation Examples Two error distributions can have the same 99% field accuracy and radically different document error rates: 1% versus 50%. This educational dataset provides reproducible, synthetic examples for understanding document AI evaluation. It helps developers test metric aggregation and explain why field accuracy alone does not describe the proportion of complete documents that need correction. How we calculated these results… See the full description on the dataset page: https://huggingface.co/datasets/pankaj9296/document-accuracy-aggregation-examples.tabular1K<n<10K0 likes26 downloads2d agoHugging Face12hoanglvuit /Legal-Documenttext100K<n<1M0 likes21 downloads10mo agoHugging Face13itsjhuang /watsonx-docs-document-type Watsonx Docs Document Type Classification This dataset is a balanced binary document-level classification subset derived from ibm-research/watsonxDocsQA. Task Classify IBM Watsonx documentation pages by their dominant user-facing purpose: conceptual: documents primarily used to understand or look up information. how-to: documents primarily used to complete a procedure or fix a problem. Splits Split conceptual how-to Total train 140 140 280… See the full description on the dataset page: https://huggingface.co/datasets/itsjhuang/watsonx-docs-document-type.texttext-classificationn<1K0 likes20 downloads5mo agoHugging Face14ClarusC64 /legal-document-version-redline-final-coherence-risk-v0.1What this dataset does You receive version history redline summary final id sent or filed id approval record mismatch flags You decide coherent or incoherent Daily use wrong attachment prevention filing version QC approval gap detection tabulartext-classificationn<1K0 likes18 downloads8mo agoHugging Face15pranay27sy /maritime_documents_tags_classificationData Understanding and Preparation Data Collection BIMCO Contracts and Clauses URL: BIMCO Contracts and Clauses Description: This website provides a wide range of standardized contracts and clauses commonly used in the maritime industry. The documents were downloaded and used as part of our dataset, offering detailed insights into industry-specific terminology and structured data. Data Description Documents Collected: A total of 217 documents were collected… See the full description on the dataset page: https://huggingface.co/datasets/pranay27sy/maritime_documents_tags_classification.texttext-classification1K<n<10K1 likes17 downloads2y agoHugging Face16abhinavdread /msme-dispute-document-corpus MSME Dispute Document Corpus (Synthetic OCR) Dataset Description This dataset contains 8,000+ synthetic document samples designed to train AI models for the Indian MSME (Micro, Small, and Medium Enterprises) dispute resolution sector. It is specifically engineered to handle Real-World OCR Noise and Adversarial Edge Cases (e.g., distinguishing a "Proforma Invoice" from a valid "Tax Invoice"). The data mimics the messy, unstructured text often found in scanned PDFs, photos… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-dispute-document-corpus.tabulartext-classification1K<n<10K0 likes17 downloads8mo agoHugging Face17abhinavdread /msme-document-presence-dataset MSME Document Presence Detection Dataset Overview This dataset is designed for training binary classification models to detect the presence of mandatory documents in MSME arbitration cases using OCR-extracted text. The dataset supports automated document completeness validation systems. Each sample represents a structured arbitration case with document-specific OCR text fields and binary presence labels. Documents Covered The dataset includes detection labels… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/msme-document-presence-dataset.tabulartext-classification10K<n<100K0 likes14 downloads8mo agoHugging Face18sriramahesh2000 /DocumentCreationtextn<1K1 likes13 downloads3y agoHugging Face19techtitans232 /medical-documentation-datasettextn<1K0 likes13 downloads2y agoHugging Face20ClarusC64 /legal-chronology-event-document-issue-coherence-risk-v0.1What this dataset does You receive timeline summary document map issue links date checks gap flags conflict flags You decide coherent or incoherent Daily use chronology QC date conflict detection missing evidence detection gap finding tabulartext-classificationn<1K0 likes13 downloads8mo agoHugging Face21hotamago /SoICT-Hackathon-2024-Legal-Document-Retrievaltext100K<n<1M0 likes12 downloads2y agoHugging Face22flaviawallen /MNLP_M3_rag_documentstext1K<n<10K0 likes12 downloads1y agoHugging Face23shearman96 /tactics-documentstext1K<n<10K0 likes11 downloads2y agoHugging Face24ushakov15 /MNLP_M2_rag_documentstext10K<n<100K0 likes11 downloads1y agoHugging Face25anasse15 /MNLP_M3_rag_documenttext10K<n<100K0 likes11 downloads1y agoHugging Face26Adityaaaa468 /Financial_Documents_datasettexttext-classification10K<n<100K2 likes11 downloads1y agoHugging Face27jason1966 /adwaittagalpallewar_medical-document-ocr-text-dataset Medical Document OCR Text Dataset Synthetic OCR-extracted text from medical documents for NLP classification Dataset Info Source: Kaggle Original Size: 22.39 MB Kaggle Downloads: 33 Files: 1 Files medical_documents_dataset.csv Mirrored from Kaggle text10K<n<100K0 likes11 downloads6mo agoHugging Face28nsjain /single-document-tokenizedtabular100K<n<1M0 likes11 downloads6mo agoHugging Face29NilsML /RETRIEVED_DOCUMENTStext100K<n<1M0 likes10 downloads1y agoHugging Face30ClarusC64 /legal-privilege-log-document-basis-waiver-risk-v0.1What this dataset does You receive doc description date author recipients privilege basis redaction choice context waiver flags You decide coherent or incoherent Daily use privilege log QC waiver risk detection disclosure challenge prep tabulartext-classificationn<1K0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.