Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface /documentation-images This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/. imagen<1K208 likes2.1m downloads11h agoHugging Face02huggingface-course /documentation-imagesimagen<1K3 likes320k downloads1y agoHugging Face03AmazonScience /document-haystack Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”. 📑 Abstract Paper The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.textquestion-answering20 likes85k downloads1y agoHugging Face04EleutherAI /wikitext_document_level Wikitext Document Level This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below. Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.text10K<n<100K18 likes84k downloads2y agoHugging Face05trl-lib /documentation-imagesimagen<1K0 likes46k downloads1mo agoHugging Face06pruna-test /documentation-mediaimagen<1K0 likes38k downloads23d agoHugging Face07jinaai /documentation-imagesimagen<1K0 likes34k downloads1y agoHugging Face08AyoubChLin /Company-document-dataset-v2 Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentdocument-question-answering100K<n<1M1 likes12k downloads11h agoHugging Face09tiiuae /documentation-imagesimagen<1K0 likes8.1k downloads1y agoHugging Face10abigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes7k downloads25d agoHugging Face11optimum /documentation-imagesThis dataset contains images used in the documentation of HuggingFace's Optimum library. imagen<1K2 likes6.1k downloads3mo agoHugging Face12mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes4k downloads27d agoHugging Face13ybelkada /documentation-imagesimagen<1K0 likes3.7k downloads3y agoHugging Face14HuggingFaceM4 /DocumentVQAimage10K<n<100K46 likes3.6k downloads3y agoHugging Face15Chunte /documentation-imagesimagen<1K0 likes3.5k downloads5mo agoHugging Face16AiAF /JFK-Assassination-Records-2025-Documents-Release0 likes3.2k downloads2y agoHugging Face17embedl /documentation-imagesimagen<1K0 likes2.6k downloads2mo agoHugging Face18incrediblecrab /scotus-case-documents Supreme Court of the United States: Case Documents The document files that the Supreme Court's website lists under Case Documents in its footer, one config per collection, each row a file with the file itself, byte for byte, and its text. Nothing here is edited by hand, and no text is corrected, normalized or generated. The pipeline and its tests are in github.com/incrediblecrab/scotus-research-service-products, and this card is rendered from the collections' manifests in the… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/scotus-case-documents.tabular1K<n<10K0 likes2.5k downloads4h agoHugging Face19CGIAR /Embrapa-ai-documents-markdown1 likes2.4k downloads3mo agoHugging Face20morzel85 /synthetic-medical-document-recognition-benchmark Synthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the dataset suitable for manual testing, product demonstrations, and workflows that… See the full description on the dataset page: https://huggingface.co/datasets/morzel85/synthetic-medical-document-recognition-benchmark.documentimage-to-text10K<n<100K1 likes2k downloads3d agoHugging Face21Voxel51 /form_understanding_in_noisy_scanned_documents_plus Dataset Card for Form Understanding in Noisy Scanned Documents Plus This is a FiftyOne dataset with 1026 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/form_understanding_in_noisy_scanned_documents_plus") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/form_understanding_in_noisy_scanned_documents_plus.imageobject-detection1K<n<10K1 likes1.9k downloads1y agoHugging Face22RoboCOIN /AI2_Alphabot_2_stamp_document AI2_Alphabot_2_stamp_document Dataset Description This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Task Preview View Video Directly Overview Total Episodes: 987 Total Frames: 369702 FPS: 30 Dataset Size: 7.18 GB Robot Name: AI2_Alphabot_2 End-Effector Type: two_finger_end_effector Teleoperation Type: vr_controller Sensors: cam_front_chest_rgb, cam_front_head_rgb, cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_stamp_document.robotics0 likes1.6k downloads3mo agoHugging Face23SherlockRamos /jurisdb-legal-documents JurisDB - Brazilian Legal Documents Dataset Dataset Description This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU). Dataset Structure . ├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/ │ ├── leis_estaduais/ │ ├── leis_federais/ │ └── ... └── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.documenttext-classificationn<1K0 likes1.6k downloads9mo agoHugging Face24RoboCOIN /Agilex_Cobot_Magic_zip_up_the_document_baggated Agilex_Cobot_Magic_zip_up_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: Agilex_Cobot_Magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pull place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.tabularrobotics100K<n<1M0 likes1.5k downloads4mo agoHugging Face25Voxel51 /document-haystack-10pages Dataset Card for document-haystack-10pages This is a FiftyOne dataset with 250 samples. It's the 10-page subset of the full dataset. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/document-haystack-10pages") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/document-haystack-10pages.imageimage-classificationn<1K1 likes1.5k downloads1y agoHugging Face26CGIAR /gardian-cigi-ai-documents-markdown1 likes1.4k downloads3mo agoHugging Face27eustlb /documentation-imagesimagen<1K0 likes1.4k downloads11mo agoHugging Face28th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M46 likes1.3k downloads2mo agoHugging Face29mannycooper /document-review-source1k New 1K title extraction corpus Independent Task0178 expansion, parsed and machine annotated in Task0180, human-reviewed under Task0186. This is NOT the 1K subset sampled from the older 4K corpus. Private research use only; original document rights are not independently verified. 1000 unique source documents: Word, Excel, PDF and PPT each 250; English 500, simplified Chinese 400, traditional Chinese 100. All 1000 human reviews collected: 711 single-title, 194 multi-title, 54… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-source1k.document1K<n<10K0 likes1.3k downloads16d agoHugging Face30CGIAR /ifpri-ai-documents-markdown GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 21,726 Total Size 623.27 MB Total Tokens 85,359,442 Total Pages 0 Languages 25 Unique Keywords 7,127 Resource Types 20 Date Generated 2026-07-31 02:55:20 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.summarization10K<n<100K1 likes1.2k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.