Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face02meta13sphere /phaseShift_shell_result_pdf Phase Resonance / IRS-DCE Topological Dynamics & Artificial Cognitive Physics Open Structural Record of Basis-Relative Reorganization in Transformer Representation Space All pdf Creative Creative Commons Attribution No Derivatives 4.0 International This repository provides comprehensive PDF research materials and Python scripts and something for mathematical proofs in the field of AI. 📢 Achievement: total 51k+ Downloads!… See the full description on the dataset page: https://huggingface.co/datasets/meta13sphere/phaseShift_shell_result_pdf.documentn<1K0 likes12k downloads10d agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes8.7k downloads2y agoHugging Face04mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes6.7k downloads2y agoHugging Face05TheFinAI /multi-ocr-source-pdfsgated MultiFinBen OCR source documents 📄 Paper · 💻 Code · 🌐 The Fin AI Raw source documents behind the OCR tasks of MultiFinBen — MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application (arXiv:2506.14028). Formerly TheFinAI/OCR_Task (the old name redirects here). This is not an evaluation set. For the curated OCR benchmarks use en-ocr, es-ocr, gr-ocr and jp-ocr. Contents Sources: U.S. SEC EDGAR filings (PDF/HTML) and… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/multi-ocr-source-pdfs.imageimage-to-text10K<n<100K1 likes3.3k downloads3d agoHugging Face06mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes2.5k downloads2y agoHugging Face07nielsr /example-pdfimagen<1K0 likes1.6k downloads3y agoHugging Face08laion /biorXiv-pdf BiorXiv Pdf BiorXiv PDF dataset is a collection of PDF documents gathered from the BiorXiv website. This initiative aims to democratize artificial intelligence research by providing researchers with access to readily available training datasets. It is part of our broader effort to publish open access research papers as collective datasets. BiorXiv is a renowned preprint publication in the field of biology and related disciplines. It is operated by Cold Spring Harbor Laboratory (CSHL)… See the full description on the dataset page: https://huggingface.co/datasets/laion/biorXiv-pdf.documentfeature-extraction1K<n<10K4 likes1.3k downloads2y agoHugging Face09dlxjj /pdf_ocrimagen<1K1 likes997 downloads1y agoHugging Face10leonardPKU /pdf_1024image1K<n<10K0 likes576 downloads2y agoHugging Face11laion /medrXiv-pdf MedrXiv Pdf Introducing MedrXiv Pdf, a dataset that offers access to all PDFs published until September 15, 2024. This resource aims to facilitate artificial intelligence research and the training of domain-specific scientific models. As part of our efforts to democratise knowledge in the scientific domain, we have compiled this dataset. While most papers included have non-restrictive and open access licences, certain PDFs may have additional restrictions. Researchers are encouraged to… See the full description on the dataset page: https://huggingface.co/datasets/laion/medrXiv-pdf.image10K<n<100K5 likes436 downloads2y agoHugging Face12philschmid /pdf-samplesimagen<1K0 likes389 downloads2y agoHugging Face13racineai /ocr-pdf-degraded OCR-PDF-Degraded Dataset Overview This dataset contains synthetically degraded document images paired with their ground truth OCR text. It addresses a critical gap in OCR model training by providing realistic document degradations that simulate real-world conditions encountered in production environments. Purpose Most OCR models are trained on relatively clean, perfectly scanned documents. However, in real-world applications, especially in the military/defense… See the full description on the dataset page: https://huggingface.co/datasets/racineai/ocr-pdf-degraded.imagetext-generation10K<n<100K3 likes325 downloads2y agoHugging Face14dreamerlin /pdf_imageimage10K<n<100K0 likes266 downloads2y agoHugging Face15TheFinAI /es-ocr-source-pdfsgated MultiFinBen Spanish OCR source documents 📄 Paper · 💻 Code · 🌐 The Fin AI Raw source documents behind the OCR tasks of MultiFinBen — MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application (arXiv:2506.14028). Formerly TheFinAI/MultiFinBen_OCR_Task (the old name redirects here). This is not an evaluation set. For the curated OCR benchmarks use en-ocr, es-ocr, gr-ocr and jp-ocr. Contents Sources: Peruvian financial… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/es-ocr-source-pdfs.documentimage-to-text10K<n<100K0 likes236 downloads3d agoHugging Face16v1v1d /pdf_images_filtered Image Dataset with Parquet Format This dataset contains images with their IDs in parquet format for efficient loading. Dataset Structure Each configuration (language) contains: image: PIL Image object id: String identifier Languages Arabic (ar): 955 images Bengali (bn): 932 images German (de): 940 images English (en): 932 images Spanish (es): 951 images French (fr): 947 images Gujarati (gu): 949 images Hindi (hi): 891 images Italian (it): 1,005 images… See the full description on the dataset page: https://huggingface.co/datasets/v1v1d/pdf_images_filtered.image10K<n<100K0 likes236 downloads8mo agoHugging Face17prithivMLmods /OpenDoc-Pdf-Preview OpenDoc-Pdf-Preview OpenDoc-Pdf-Preview is a compact visual preview dataset containing 6,000 high-resolution document images extracted from PDFs. This dataset is designed for Image-to-Text tasks such as document OCR pretraining, layout understanding, and multimodal document analysis. Dataset Summary Modality: Image-to-Text Content Type: PDF-based document previews Number of Samples: 6,000 Language: English Format: Parquet Split: train only Size: 606 MB License: Apache… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDoc-Pdf-Preview.documentimage-to-text1K<n<10K2 likes233 downloads1y agoHugging Face18yang-syzng /GS_PDF_repoimage1K<n<10K0 likes213 downloads3mo agoHugging Face19Cyber-security-final-project /HARMLESS_Synthetic_Injected_PDFs_EDA Injected PDFs - EDA and Evaluation Corpus This repository holds the exploratory data analysis for a project on detecting harmless-but-real attack payloads injected into PDF files, together with the dataset that analysis produced. The project has two halves, both in the notebook Final_project_V7_EDA.ipynb: Question Input Part 1 Is our synthetic corpus a stand-in for real malware, or is it something else? The published CIC feature table (11,126 x 34) Part 2 Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.imagetext-classification1K<n<10K0 likes166 downloads2mo agoHugging Face20dauvannam321 /fineweb-vi-pdf-layout Fineweb VI PDF Layout (v2) Vietnamese web pages (from HuggingFaceFW/fineweb-2) rendered to PDF, with per-page layout annotations for VLM training. Each row = 1 document: Column Type Description doc_name string Document ID (e.g. doc_000) url string Source URL (metadata) pdf binary Rendered document.pdf layouts_merged string All per-page layout JSONs merged: {"num_pages": N, "pages": [...]} pages image list Rendered page images (page_*.png) source_html string… See the full description on the dataset page: https://huggingface.co/datasets/dauvannam321/fineweb-vi-pdf-layout.imageimage-to-text1K<n<10K0 likes160 downloads19d agoHugging Face21sh61309 /pecan-pdfsdocumentn<1K0 likes134 downloads7mo agoHugging Face22Misraj /KITAB_pdf_to_markdown_reviewed KITAB_pdf_to_markdown_reviewed (Corrected KITAB-Bench PDF→Markdown) Short description. A carefully reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. We fixed ground-truth errors (hallucinated text, missing page numbers, omissions of small-font text) and standardized formatting to provide a reliable benchmark for model comparison. TL;DR ✅ Human-verified ground truth for Arabic PDF→Markdown ✅ Removes hallucinations and fills… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/KITAB_pdf_to_markdown_reviewed.imagen<1K4 likes124 downloads1y agoHugging Face23gigant /pdfvqa Dataset Card for "pdfvqa" More Information needed image10K<n<100K6 likes113 downloads2y agoHugging Face24laion /chemrXiv-pdf ChemrXiv Pdf Introducing ChemrXiv Pdf, a dataset that offers access to all PDFs published until September 15, 2024. This resource aims to facilitate artificial intelligence research and the training of domain-specific scientific models. As part of our efforts to democratize knowledge in the scientific domain, we have compiled this dataset. While most papers included have non-restrictive and open access licenses, certain PDFs may have additional restrictions. Researchers are encouraged… See the full description on the dataset page: https://huggingface.co/datasets/laion/chemrXiv-pdf.documentsummarization1K<n<10K0 likes88 downloads2y agoHugging Face25PDFPages /colpali_train_set_pdf_subset_sample_with_labelsimage1K<n<10K0 likes83 downloads2y agoHugging Face26nirantk /finance-pdf-vqaimage1K<n<10K4 likes82 downloads2y agoHugging Face27lhoestq /resumes-raw-pdf-for-ocrExtracted lists of pages from PDF resumes and the PDF texts. Created using this code: import io import PIL.Image from datasets import load_dataset def render(pdf): images = [] for page in pdf.pages: buffer = io.BytesIO() page.to_image(height=840).save(buffer) images.append(PIL.Image.open(buffer)) return images def extract_text(pdf): return "\n".join(page.extract_text() for page in pdf.pages) ds = load_dataset("d4rk3r/resumes-raw-pdf", split="train")… See the full description on the dataset page: https://huggingface.co/datasets/lhoestq/resumes-raw-pdf-for-ocr.image1K<n<10K3 likes82 downloads2y agoHugging Face28dauvannam321 /fineweb-pdf-vieimage1K<n<10K0 likes80 downloads21d agoHugging Face29rinabuoy /khmer-doc-markdown-pdf-kh3image1K<n<10K0 likes69 downloads2mo agoHugging Face30Reza2kn /tables-pdf-20260823-pagesimagen<1K1 likes61 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.