Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Oitaaa /GDP.pdf GDP.pdf GDP.pdf GDP.pdf measures professional multimodal reasoning over the documents the economy actually runs on: dense reports, contracts, filings, and records in the messy real-world formats professionals work from. It was cited in Anthropic's Fable 5 and Mythos 5 model card. What it tests The benchmark contains 100 real-world prompts and PDFs pulled directly from professional workflows across ten domains: Finance, Healthcare, Legal, STEM/Research, Engineering, Construction… See the full description on the dataset page: https://huggingface.co/datasets/Oitaaa/GDP.pdf.documentdocument-question-answeringn<1K0 likes25k downloads3mo agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face03RevolutionCrossroads /nara_revolutionary_war_pension_files_PDFs Dataset Card for American Revolutionary War Pension Files - File-Level Dataset Summary A dataset derived from the National Archives and Records Administration (NARA) series Case Files of Pension and Bounty-Land Warrant Applications Based on American Revolutionary War Service (NARA Catalog Series, NAID 300022). This dataset provides a file-level representation of Revolutionary War pension records, aggregating individual page records into complete pension files… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/nara_revolutionary_war_pension_files_PDFs.documentimage-to-text10K<n<100K0 likes13k downloads5d agoHugging Face04mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes8.7k downloads2y agoHugging Face05pdfqa /pdfQA-Benchmark pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval-augmented QA End-to-end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Benchmark.documentquestion-answering5 likes7k downloads7mo agoHugging Face06mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes6.7k downloads2y agoHugging Face07AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes5.4k downloads1y agoHugging Face08pixparse /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/pdfa-eng-wds.textimage-to-text1K<n<10K161 likes5k downloads3y agoHugging Face09Edinburgh-Claire /pdfQA-Benchmark pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval-augmented QA End-to-end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/Edinburgh-Claire/pdfQA-Benchmark.documentquestion-answering0 likes3.5k downloads3mo agoHugging Face10ksokolovic /dec-pdfsdocumentn<1K0 likes2.5k downloads1y agoHugging Face11mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes2.5k downloads2y agoHugging Face12laion /CS-Arxiv-PDFs-08-25text10M<n<100M2 likes2.3k downloads1y agoHugging Face13BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes1.9k downloads10mo agoHugging Face14KaiserML /Techie_Raw_PDFtabular100K<n<1M0 likes1.8k downloads3y agoHugging Face15ranWang /UN_Historical_PDF_Article_Text_Corpus python dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="train") or dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="randomTest") lang_list = ["ar", "en", "es", "fr", "ru", "zh"] for row in dataset: # 获取pdf文章内容 for lang in lang_list: # type == str lang_match_file_content = row[lang] # 如果按页分割 lang_match_file_pages_content = lang_match_file_content.split("\n----\n") text100K<n<1M2 likes1.6k downloads3y agoHugging Face16nakasyou /eiken-pdfdocumentn<1K0 likes1.5k downloads10mo agoHugging Face17sailor2 /sea-pdf-texttext10M<n<100M1 likes1.3k downloads2y agoHugging Face18LLMDH /marianne_pdf_7text10K<n<100K0 likes1.3k downloads2y agoHugging Face19LLMDH /marianne_pdf_5text100K<n<1M0 likes1.1k downloads2y agoHugging Face20PleIAs /WTO-PDFtext100K<n<1M8 likes1k downloads2y agoHugging Face21LLMDH /marianne_pdf_9text100K<n<1M0 likes996 downloads2y agoHugging Face22LLMDH /marianne_pdf_3text100K<n<1M0 likes975 downloads2y agoHugging Face23jonathanli /pdfa-eng-wds-filteredtext10M<n<100M0 likes947 downloads5mo agoHugging Face24AtlasUnified /atlas-pdf-img-cluster Atlas PDF Image Cluster Dataset Derives from the following Python Pipeline code: https://github.com/atlasunified/PDF-to-Image-Cluster Dataset Description This dataset is a collection of text extracted from PDF files, originating from various online resources. The dataset was generated using a series of Python scripts forming a robust pipeline that automated the tasks of downloading, converting, and managing the data. Dataset Summary Sample JPG Corresponding… See the full description on the dataset page: https://huggingface.co/datasets/AtlasUnified/atlas-pdf-img-cluster.textimage-classification10M<n<100M2 likes814 downloads3y agoHugging Face25LLMDH /marianne_pdf_4text10K<n<100K0 likes795 downloads2y agoHugging Face26PleIAs /AMF-PDFtext100K<n<1M7 likes772 downloads2y agoHugging Face27tolo6474 /GDP.pdf GDP.pdf GDP.pdf GDP.pdf measures professional multimodal reasoning over the documents the economy actually runs on: dense reports, contracts, filings, and records in the messy real-world formats professionals work from. It was cited in Anthropic's Fable 5 and Mythos 5 model card. What it tests The benchmark contains 100 real-world prompts and PDFs pulled directly from professional workflows across ten domains: Finance, Healthcare, Legal, STEM/Research, Engineering, Construction… See the full description on the dataset page: https://huggingface.co/datasets/tolo6474/GDP.pdf.documentdocument-question-answeringn<1K0 likes725 downloads3mo agoHugging Face28LLMDH /hal_pdf_extratext10K<n<100K0 likes673 downloads2y agoHugging Face29LLMDH /marianne_pdf_8text100K<n<1M0 likes672 downloads2y agoHugging Face30adorkin /olmocr_science_pdfs-literaturehttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-literature text1M<n<10M0 likes650 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.