Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface /documentation-images This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/. imagen<1K208 likes2.1m downloads2d agoHugging Face02huggingface-course /documentation-imagesimagen<1K3 likes315k downloads1y agoHugging Face03EleutherAI /wikitext_document_level Wikitext Document Level This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below. Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.text10K<n<100K18 likes87k downloads2y agoHugging Face04AmazonScience /document-haystack Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”. 📑 Abstract Paper The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.textquestion-answering20 likes82k downloads1y agoHugging Face05trl-lib /documentation-imagesimagen<1K0 likes45k downloads1mo agoHugging Face06pruna-test /documentation-mediaimagen<1K0 likes37k downloads26d agoHugging Face07jinaai /documentation-imagesimagen<1K0 likes34k downloads1y agoHugging Face08AyoubChLin /Company-document-dataset-v2 Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentdocument-question-answering100K<n<1M2 likes12k downloads3d agoHugging Face09tiiuae /documentation-imagesimagen<1K0 likes8.1k downloads1y agoHugging Face10abigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes6.5k downloads27d agoHugging Face11optimum /documentation-imagesThis dataset contains images used in the documentation of HuggingFace's Optimum library. imagen<1K2 likes6k downloads3mo agoHugging Face12HuggingFaceM4 /DocumentVQAimage10K<n<100K46 likes4.2k downloads3y agoHugging Face13ybelkada /documentation-imagesimagen<1K0 likes3.8k downloads3y agoHugging Face14Chunte /documentation-imagesimagen<1K0 likes3.5k downloads5mo agoHugging Face15mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3.5k downloads1mo agoHugging Face16incrediblecrab /scotus-case-documents Supreme Court of the United States: Case Documents The document files that the Supreme Court's website lists under Case Documents in its footer, one config per collection, each row a file with the file itself, byte for byte, and its text. Nothing here is edited by hand, and no text is corrected, normalized or generated. The pipeline and its tests are in github.com/incrediblecrab/scotus-research-service-products, and this card is rendered from the collections' manifests in the… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/scotus-case-documents.tabular1K<n<10K0 likes3.1k downloads8h agoHugging Face17AiAF /JFK-Assassination-Records-2025-Documents-Release0 likes3.1k downloads2y agoHugging Face18embedl /documentation-imagesimagen<1K0 likes2.5k downloads2mo agoHugging Face19CGIAR /Embrapa-ai-documents-markdown1 likes2.3k downloads3mo agoHugging Face20SherlockRamos /jurisdb-legal-documents JurisDB - Brazilian Legal Documents Dataset Dataset Description This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU). Dataset Structure . ├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/ │ ├── leis_estaduais/ │ ├── leis_federais/ │ └── ... └── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.documenttext-classificationn<1K0 likes2.2k downloads9mo agoHugging Face21morzel85 /synthetic-medical-document-recognition-benchmark Synthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the dataset suitable for manual testing, product demonstrations, and workflows that… See the full description on the dataset page: https://huggingface.co/datasets/morzel85/synthetic-medical-document-recognition-benchmark.documentimage-to-text10K<n<100K1 likes2.1k downloads6d agoHugging Face22Voxel51 /form_understanding_in_noisy_scanned_documents_plus Dataset Card for Form Understanding in Noisy Scanned Documents Plus This is a FiftyOne dataset with 1026 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/form_understanding_in_noisy_scanned_documents_plus") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/form_understanding_in_noisy_scanned_documents_plus.imageobject-detection1K<n<10K1 likes1.7k downloads1y agoHugging Face23RoboCOIN /Agilex_Cobot_Magic_zip_up_the_document_baggated Agilex_Cobot_Magic_zip_up_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: Agilex_Cobot_Magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pull place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.tabularrobotics100K<n<1M0 likes1.5k downloads4mo agoHugging Face24CGIAR /gardian-cigi-ai-documents-markdown1 likes1.4k downloads3mo agoHugging Face25Voxel51 /document-haystack-10pages Dataset Card for document-haystack-10pages This is a FiftyOne dataset with 250 samples. It's the 10-page subset of the full dataset. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/document-haystack-10pages") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/document-haystack-10pages.imageimage-classificationn<1K1 likes1.4k downloads1y agoHugging Face26eustlb /documentation-imagesimagen<1K0 likes1.4k downloads11mo agoHugging Face27HumynLabs /French_Documents_Dataset_PDF French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.documentn<1K0 likes1.4k downloads11mo agoHugging Face28th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M46 likes1.3k downloads2mo agoHugging Face29mannycooper /document-review-source1k New 1K title extraction corpus Independent Task0178 expansion, parsed and machine annotated in Task0180, human-reviewed under Task0186. This is NOT the 1K subset sampled from the older 4K corpus. Private research use only; original document rights are not independently verified. 1000 unique source documents: Word, Excel, PDF and PPT each 250; English 500, simplified Chinese 400, traditional Chinese 100. All 1000 human reviews collected: 711 single-title, 194 multi-title, 54… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-source1k.document1K<n<10K0 likes1.3k downloads18d agoHugging Face30HumynLabs /Arabic_Documents_Dataset_PDF Arabic Documents Dataset (PDF) This dataset contains a collection of Arabic-language documents in PDF format. The corpus includes books, articles, reports, and educational materials written in Modern Standard Arabic and regional variants. It is curated to support AI research in document understanding, Arabic OCR, and text extraction from complex layouts. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Arabic_Documents_Dataset_PDF.documentn<1K0 likes1.2k downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.