Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /wikitext_document_level Wikitext Document Level This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below. Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.text10K<n<100K18 likes87k downloads2y agoHugging Face02AmazonScience /document-haystack Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”. 📑 Abstract Paper The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.textquestion-answering20 likes82k downloads1y agoHugging Face03AyoubChLin /Company-document-dataset-v2 Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentdocument-question-answering100K<n<1M2 likes12k downloads3d agoHugging Face04HuggingFaceM4 /DocumentVQAimage10K<n<100K46 likes4.2k downloads3y agoHugging Face05mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3.5k downloads1mo agoHugging Face06incrediblecrab /scotus-case-documents Supreme Court of the United States: Case Documents The document files that the Supreme Court's website lists under Case Documents in its footer, one config per collection, each row a file with the file itself, byte for byte, and its text. Nothing here is edited by hand, and no text is corrected, normalized or generated. The pipeline and its tests are in github.com/incrediblecrab/scotus-research-service-products, and this card is rendered from the collections' manifests in the… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/scotus-case-documents.tabular1K<n<10K0 likes3.1k downloads12h agoHugging Face07th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M46 likes1.3k downloads2mo agoHugging Face08mannycooper /document-review-source1k New 1K title extraction corpus Independent Task0178 expansion, parsed and machine annotated in Task0180, human-reviewed under Task0186. This is NOT the 1K subset sampled from the older 4K corpus. Private research use only; original document rights are not independently verified. 1000 unique source documents: Word, Excel, PDF and PPT each 250; English 500, simplified Chinese 400, traditional Chinese 100. All 1000 human reviews collected: 711 single-title, 194 multi-title, 54… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-source1k.document1K<n<10K0 likes1.3k downloads18d agoHugging Face09jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes974 downloads1y agoHugging Face10chen-doc-ai /Unified_Document_Understanding_Dataset Dataset Card for Read-Parsing-Describe: Unified Scientific Document Understanding Read-Parsing-Describe (Unified Scientific Document Understanding) is a pioneering multimodal benchmark designed to train and evaluate models on the complex structures of scientific documents. Unlike traditional document datasets that treat visual elements merely as isolated layout blocks, RPD transforms document parsing into an accessibility-driven, cross-modal grounding task. 🌟 Key… See the full description on the dataset page: https://huggingface.co/datasets/chen-doc-ai/Unified_Document_Understanding_Dataset.imageimage-to-text10K<n<100K1 likes851 downloads12d agoHugging Face11ankitjh4 /bharat-government-documents Bharat Guide: screened Indian public information documents This snapshot contains 64,964 distinct normalized text bodies and 77,526 source records. Generated 2026-10-03T01:10:02.179230+00:00. Contents and provenance One Parquet row represents one normalized source body, with original extracted text, source title, URL, extraction quality, observed retrieval time, rule version, topic labels and all current-body source aliases. Duplicate URLs/editions are preserved… See the full description on the dataset page: https://huggingface.co/datasets/ankitjh4/bharat-government-documents.textquestion-answering10K<n<100K59 likes782 downloads6d agoHugging Face12singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes694 downloads11mo agoHugging Face13himalaya-ai /ocr-document-processing-eval ocr_document_processing_eval Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks. Repo: himalaya-ai/ocr-document-processing-eval Task: document_processing_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.imageimage-to-textn<1K1 likes676 downloads4mo agoHugging Face14hf-internal-testing /document-visual-retrieval-test Model Card: Document Visual Retrieval Test (internal) Dataset Overview This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.imagen<1K1 likes675 downloads2y agoHugging Face15ZamAI-Pashto /zamai-pashto-documents ZamAI-Pashto Documents This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language. Project Structure data/: Contains scanned documents, extracted text, translations, and summaries. annotations/: OCR bounding boxes, handwriting labels, and domain tags. scripts/: OCR processing, text cleaning, and translation alignment scripts. configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes663 downloads2mo agoHugging Face16hotchpotch /cc100-ja-documents cc100-ja-documents HuggingFace で公開されている cc100 / cc100-ja は line 単位の分割のため、document 単位に結合したものです。 ライセンスはオリジナルのcc100 に準拠します。 text10M<n<100M4 likes634 downloads2y agoHugging Face17nyuuzyou /znanio-documents Dataset Card for Znanio.ru Educational Documents Dataset Summary This dataset contains 588,545 educational document files from the znanio.ru platform, a resource for teachers, educators, students, and parents providing diverse educational content. Znanio.ru has been a pioneer in educational technologies and distance learning in the Russian-speaking internet since 2009. The dataset includes a small portion of English language content, primarily for language learning… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-documents.texttext-classification100K<n<1M0 likes611 downloads2y agoHugging Face18Samoed /msmarco-documenttext1M<n<10M0 likes582 downloads11mo agoHugging Face19eddmpython /cleangov-local-documents 지방자치단체 예산서·결산서·재정공시·계약현황·세입세출 원문 색인 (지방재정365 "우리 지자체" 링크 API) 과 원문 각 자치단체가 자기 홈페이지에 공개하는 예산서, 결산서(성과보고서·성인지·기금 결산 포함), 재정공시, 계약현황, 세입세출현황 원문의 자치단체별 연도별 URL 색인과 그 원문. 색인은 지방재정365 허브의 "우리 지자체 *" 5 서비스(SETLK 결산서·BUDLK 예산서·FINLK 재정공시·CONLK 계약현황·REVLK 세입세출현황)가 회계연도 × 243 자치단체 × 링크(lnk_url_nm) 로 준다. 원문 자체는 자치단체 서식(PDF·HWP·HWPX·XLSX)이고 세부 산출내역과 판단에 필요한 문맥이 여기 있다. 지방 성과보고서(EVAL-004)와 자체감사·결산검사(AUDIT-003) 도 이 색인으로 각 자치단체의 결산 문서에 도달한다. 출처: https://www.lofin365.go.kr/portal/LF5100000.do 이용… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-documents.text10K<n<100K0 likes520 downloads6d agoHugging Face20Lukaszl /clearocr-invoice-document-ai clearOCR Invoice Document AI Dataset This dataset shows a complete invoice document AI workflow built around clearOCR. It contains 423 high-confidence invoice examples with: original invoice images, OCR text generated by clearOCR, Markdown reconstruction of the document, structured invoice JSON generated by a local fine-tuned extraction model, visual verification metadata. The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.imageimage-to-textn<1K0 likes509 downloads5mo agoHugging Face21biglam /icdar2021-historical-document-dating ICDAR 2021 Historical Document Classification — Task 2 (Dating) 13,810 manuscript page images labelled with the period in which they were produced. Images come from e-codices, the virtual manuscript library of Switzerland. Split Images Date range Median span Dated to a single year train 11,294 800–1899 45 years 1,409 test 2,516 800–1921 49 years 264 The label is an interval, not a year Palaeographers date a manuscript to a range, and the width of… See the full description on the dataset page: https://huggingface.co/datasets/biglam/icdar2021-historical-document-dating.imageimage-classification10K<n<100K2 likes451 downloads2mo agoHugging Face22KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes411 downloads4mo agoHugging Face23dtunkelang /bag-of-documents Bag-of-Documents: Product Search Dataset Blog post: Distilling Retrieval Pipelines to a Single Embedding Model Live demo: huggingface.co/spaces/dtunkelang/bag-of-documents-demo Code: github.com/dtunkelang/bag-of-documents Dataset Description A large-scale bag-of-documents dataset for e-commerce product search, built on Amazon product data. Each search query is represented as a distribution of relevant products in embedding space, captured by a centroid vector and a… See the full description on the dataset page: https://huggingface.co/datasets/dtunkelang/bag-of-documents.tabularsentence-similarityn<1K5 likes405 downloads5mo agoHugging Face24hulk10 /conseil-administratives-appel-full-documents Décisions de Justice Administrative Françaises Description du dataset Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.texttext-classification1K<n<10K1 likes399 downloads9d agoHugging Face25hulk10 /legi-full-documentstext100K<n<1M0 likes384 downloads9d agoHugging Face26prg-unibe /dodis-historical-documents Dataset Overview The Swiss Historical Archive of DOdis Works (SHADOW) is a large-scale benchmark dataset derived from the resources of Dodis, an independent research center dedicated to the study of Swiss foreign policy and Switzerland’s international relations. Dodis has curated and published nearly 50,000 key documents that document the administrative practices and decision-making processes of the Swiss federal administration. The underlying corpus consists of notes, letters… See the full description on the dataset page: https://huggingface.co/datasets/prg-unibe/dodis-historical-documents.tabularsummarization10K<n<100K5 likes376 downloads4mo agoHugging Face27agentlans /en-document-classification English Document Classification Dataset This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. Dataset Summary The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.texttext-classification1M<n<10M1 likes352 downloads1mo agoHugging Face28hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes349 downloads9d agoHugging Face29hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes344 downloads9d agoHugging Face30LegionIntel /named_entity_recognition_document_contexttabular1M<n<10M9 likes335 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.