Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Chelsea707 /arxiv-cs-2020-2025-pdfs24 likes454k downloads10mo agoHugging Face02gutoportelaa /dom-pi-pdfs-2025 DOM-PI 2025 — PDFs-fonte (Diário Oficial dos Municípios do Piauí) PDFs originais das publicações de 2025 do Diário Oficial dos Municípios do Piauí, organizados por Território de Desenvolvimento. São a fonte da qual o corpus textual foi extraído por OCR/parsing. 41.617 PDFs · ~70 GB. Dataset de texto derivado (carregável, com limpeza e tiers de qualidade): gutoportelaa/dom-pi-corpus-2025. Cobertura de PDFs (parcial): presentes 7 territórios — tabuleiros_alto_parnaiba… See the full description on the dataset page: https://huggingface.co/datasets/gutoportelaa/dom-pi-pdfs-2025.document10K<n<100K1 likes136k downloads4mo agoHugging Face03atom-in-the-universe /cc-pdf-links1 likes56k downloads3y agoHugging Face04JBrightmanAI /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/pdfa-eng-wds.image-to-text10M<n<100M0 likes47k downloads3mo agoHugging Face05quejing /20000-meishi-pdf3 likes29k downloads4mo agoHugging Face06Oitaaa /GDP.pdf GDP.pdf GDP.pdf GDP.pdf measures professional multimodal reasoning over the documents the economy actually runs on: dense reports, contracts, filings, and records in the messy real-world formats professionals work from. It was cited in Anthropic's Fable 5 and Mythos 5 model card. What it tests The benchmark contains 100 real-world prompts and PDFs pulled directly from professional workflows across ten domains: Finance, Healthcare, Legal, STEM/Research, Engineering, Construction… See the full description on the dataset page: https://huggingface.co/datasets/Oitaaa/GDP.pdf.documentdocument-question-answeringn<1K0 likes19k downloads3mo agoHugging Face07ayushsinghal1510 /msf-pdfsdocumentn<1K0 likes18k downloads10mo agoHugging Face08mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face09RevolutionCrossroads /nara_revolutionary_war_pension_files_PDFs Dataset Card for American Revolutionary War Pension Files - File-Level Dataset Summary A dataset derived from the National Archives and Records Administration (NARA) series Case Files of Pension and Bounty-Land Warrant Applications Based on American Revolutionary War Service (NARA Catalog Series, NAID 300022). This dataset provides a file-level representation of Revolutionary War pension records, aggregating individual page records into complete pension files… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/nara_revolutionary_war_pension_files_PDFs.documentimage-to-text10K<n<100K0 likes16k downloads12h agoHugging Face10mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes15k downloads2y agoHugging Face11mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes13k downloads2y agoHugging Face12mlfoundations /MINT-1T-PDF-CC-2023-40 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-40.image-to-text100B<n<1T9 likes12k downloads2y agoHugging Face13mlfoundations /MINT-1T-PDF-CC-2023-06 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-06.image-to-text100B<n<1T10 likes11k downloads2y agoHugging Face14mlfoundations /MINT-1T-PDF-CC-2024-18 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-18.image-to-text100B<n<1T31 likes11k downloads2y agoHugging Face15meta13sphere /phaseShift_shell_result_pdf Phase Resonance / IRS-DCE Topological Dynamics & Artificial Cognitive Physics Open Structural Record of Basis-Relative Reorganization in Transformer Representation Space All pdf Creative Creative Commons Attribution No Derivatives 4.0 International This repository provides comprehensive PDF research materials and Python scripts and something for mathematical proofs in the field of AI. 📢 Achievement: total 51k+ Downloads!… See the full description on the dataset page: https://huggingface.co/datasets/meta13sphere/phaseShift_shell_result_pdf.documentn<1K0 likes8.9k downloads5d agoHugging Face16pixparse /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/pdfa-eng-wds.textimage-to-text1K<n<10K161 likes7.4k downloads3y agoHugging Face17pdfqa /pdfQA-Benchmark pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval-augmented QA End-to-end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Benchmark.documentquestion-answering5 likes7k downloads7mo agoHugging Face18LLM4SCIENCE /uparxive_boxed_pdfimage-to-text100M<n<1B0 likes6.6k downloads2y agoHugging Face19redmoddata /dolmino_olmocr_pdfs0 likes6.5k downloads9mo agoHugging Face20MLP-BAI /bai-pdfsdocument10K<n<100K2 likes6.5k downloads1y agoHugging Face21AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes6.4k downloads1y agoHugging Face22topdu /output_pdf_lmdb0 likes5.9k downloads1y agoHugging Face23di-zhang-fdu /libretexts_pdfsdocument1K<n<10K1 likes5.8k downloads1y agoHugging Face24forcemultiplier /supreme_court_opinions_corpus_pdfwebAug240 likes4.2k downloads2y agoHugging Face25eltokh7 /shorouk-pdf-archive Shorouk PDF archive Original newspaper PDFs from Al Shorouk's public archive. Filenames use YYYY-MM-DD.pdf. As verified in this update, the dataset contains 6,276 PDFs, dated 2009-02-01 through 2026-09-08. The September 8 edition was already available from the publisher when checked on September 7 in New York. Coverage and remaining gaps This update added 518 issues and filled all calendar gaps from January 1, 2025 through September 8, 2026. 153 earlier dates… See the full description on the dataset page: https://huggingface.co/datasets/eltokh7/shorouk-pdf-archive.document1K<n<10K0 likes3.4k downloads28d agoHugging Face26Edinburgh-Claire /pdfQA-Benchmark pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval-augmented QA End-to-end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/Edinburgh-Claire/pdfQA-Benchmark.documentquestion-answering0 likes3.4k downloads2mo agoHugging Face27mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes3.2k downloads2y agoHugging Face28LLMDH /marianne_pdf_7text10K<n<100K0 likes2.8k downloads2y agoHugging Face29siyrus /BToks-vidore_rag_pdf BToks ViDoRe PDF This dataset repository contains Lance-format converted data used by the open-source reproduction code for Bottleneck Tokens for Unified Multimodal Retrieval (arXiv:2604.11095). Source Converted from vidore/colpali_train_set. Subset/view: pdf. This repository does not change upstream ownership, licensing, citation requirements, or usage restrictions. Format The data is stored as Lance tables for the BToks/VLM2Emb training and… See the full description on the dataset page: https://huggingface.co/datasets/siyrus/BToks-vidore_rag_pdf.image-to-text0 likes2.6k downloads4mo agoHugging Face30ksokolovic /dec-pdfsdocumentn<1K0 likes2.5k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.