Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KaiserML /Techie_Raw_PDFtabular100K<n<1M0 likes2k downloads3y agoHugging Face02BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes1.8k downloads9mo agoHugging Face03nakasyou /eiken-pdfdocumentn<1K0 likes1.6k downloads10mo agoHugging Face04vasilisplavos /unboxgov-khmdhs-pdfs KHMDHS Contract Attachment PDFs Per-contract attachment PDFs from the Greek public-procurement registry KHMDHS, downsampled with Ghostscript /ebook (~150 DPI) for fast viewing and compact storage. Files are packed into per-month SQLite shards (pdfs/<YYYY-MM>.sqlite, one row per ΑΔΑΜ, keyed by referenceNumber). These are lossy, downsampled copies. The authoritative originals remain at KHMDHS: https://cerpp.eprocurement.gov.gr/khmdhs-opendata/contract/attachment/{ΑΔΑΜ}. This… See the full description on the dataset page: https://huggingface.co/datasets/vasilisplavos/unboxgov-khmdhs-pdfs.tabular100K<n<1M0 likes501 downloads22d agoHugging Face05mlfoundations-dev /pdf_science_questions_verified_r1_traces__2_24_25 Dataset card for pdf_science_questions_verified_r1_traces__2_24_25 This dataset was made with Curator. Dataset details A sample from the dataset: { "url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "success": true, "page_count": 37, "page_number": 1, "question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verified_r1_traces__2_24_25.tabular1K<n<10K0 likes441 downloads2y agoHugging Face06Wikit /pdf-parsing-bench-resultstabular100K<n<1M0 likes422 downloads2y agoHugging Face07nickua /ICLR-pdfs Dataset Card for "ICLR-pdfs" More Information needed tabular10K<n<100K0 likes359 downloads3y agoHugging Face08mlfoundations-dev /PDF_and_SCP_unfiltered_organic_chemistry_questionstabular10K<n<100K0 likes284 downloads1y agoHugging Face09tasal9 /Pashto-Textbooks-PDFs-Corpus Pashto Textbooks and PDFs Corpus Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tabulartext-generationn<1K0 likes182 downloads2mo agoHugging Face10Cyber-security-final-project /HARMLESS_Synthetic_Injected_PDFs_EDA Injected PDFs - EDA and Evaluation Corpus This repository holds the exploratory data analysis for a project on detecting harmless-but-real attack payloads injected into PDF files, together with the dataset that analysis produced. The project has two halves, both in the notebook Final_project_V7_EDA.ipynb: Question Input Part 1 Is our synthetic corpus a stand-in for real malware, or is it something else? The published CIC feature table (11,126 x 34) Part 2 Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.imagetext-classification1K<n<10K0 likes165 downloads2mo agoHugging Face11mlfoundations-dev /pdf_science_questions_verifiable_r1_traces__2_24_25 Dataset card for pdf_science_questions_verifiable_r1_traces__2_24_25 This dataset was made with Curator. Dataset details A sample from the dataset: { "url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "success": true, "page_count": 37, "page_number": 1, "question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verifiable_r1_traces__2_24_25.tabular1K<n<10K0 likes156 downloads2y agoHugging Face12nickua /ICLR-pdfs-linebreaks Dataset Card for "ICLR-pdfs-linebreaks" More Information needed tabular10K<n<100K0 likes123 downloads3y agoHugging Face13Cyber-security-final-project /Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition Injected PDFs - Model Evaluation This repository holds the model evaluation stage of a project on detecting harmless-but-real attack payloads injected into PDF files, together with the artefacts it produced for the application. Nothing is trained here. Seven off-the-shelf models are measured against the same 1,100 PDFs, and the two winners are exported for the app to load. Question Candidates Winner Part A Which files look like this one? 3 embedding models x 2 inputs… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition.tabulartext-classification1K<n<10K0 likes107 downloads2mo agoHugging Face14StentorLabs /StenCore-PDF StenCore — FinePDFs-Edu Curated By StentorLabs StenCore is StentorLabs' first dataset release — a quality-first resource built from the ground up for training language models. Derived from HuggingFaceFW/finepdfs-edu, every document passed a 14-stage automated curation pipeline including heuristic filtering, language ID, PII redaction, toxicity screening, eval decontamination, multi-strategy deduplication, KenLM + neural perplexity scoring, and domain reweighting. ⚠️ Privacy &… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/StenCore-PDF.tabulartext-generation100K<n<1M1 likes103 downloads7mo agoHugging Face15mlfoundations-dev /pdf_qa_r1_annotated_verifiedtabular100K<n<1M0 likes100 downloads2y agoHugging Face16taqbaylit /ayamun-pdfs230 pdf files from Ayamun. documentn<1K1 likes86 downloads5mo agoHugging Face17manj0220 /pdf-steganalysis-corpus PDF Steganalysis Corpus v2 A research benchmark for clean-versus-stego detection and method attribution, plus a separate downloadable demonstration made from original synthetic PDFs. "Clean" means the original document before our embedding, not malware-free. This is not a malware collection. What to download Files Purpose historical-{train,validation,test}.parquet 13,637 verified historical stego records + 1,939 clean originals, with labels and 38… See the full description on the dataset page: https://huggingface.co/datasets/manj0220/pdf-steganalysis-corpus.document10K<n<100K0 likes71 downloads14d agoHugging Face18cminst /imslp-pdf-index IMSLP PDF Index This dataset is the canonical PDF-level index for the ReScore IMSLP PDF collection. It contains one row per unique IMSLP PDF and points to the PDF payload stored in cminst/imslp-raw-pdf-collection. The PDF payload repository is append-only and may contain duplicate rows from retry launches. This index is deduplicated by imslp_id; duplicate content was validated to have identical SHA256, byte size, and page count before publishing. Summary Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cminst/imslp-pdf-index.tabularimage-to-text10K<n<100K0 likes57 downloads2mo agoHugging Face19C10X /pdftabular100K<n<1M0 likes55 downloads11mo agoHugging Face20mlfoundations-dev /all_filtered_unverified_pdfs_pipelinetabular10K<n<100K0 likes54 downloads2y agoHugging Face21mlfoundations-dev /organic_chemistry_pdftabularn<1K0 likes54 downloads2y agoHugging Face22mlfoundations-dev /organic_chemistry_pdf_word_searchtabular10K<n<100K1 likes54 downloads2y agoHugging Face23CentificAIResearch /Healthcare.pdf Healthcare.pdf — representative release (v1.0) A PDF-grounding benchmark for healthcare document work: 25 expert-authored tasks grounded in 25 real healthcare documents, covering 7 occupations across clinical and pharmacy practice. Each task puts a practitioner in a realistic situation, gives them a real document, and asks a sequence of sub-questions that must be answered from that document. Answers are graded against a four-tier rubric. Tasks 25 Source documents… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Healthcare.pdf.documentquestion-answeringn<1K0 likes47 downloads11d agoHugging Face24ranWang /UN_PDF_RECORD_SET Dataset Card for "UN_PDF_RECORD_SET" More Information needed tabular1M<n<10M0 likes38 downloads3y agoHugging Face25EasyPDF /pdf-upload-caps-2026 PDF upload caps on public portals (2026) How large can a PDF be before a government, university or job portal rejects it? This dataset records the published file size limit of 162 portals in France, the United States, the United Kingdom, Germany, Spain and India, each with the exact wording of the limit and a link to the official page where it was found. It was collected in September 2026 for the EasyPDF study The 1 MB Problem: PDF File Size Statistics for 2026 (French version:… See the full description on the dataset page: https://huggingface.co/datasets/EasyPDF/pdf-upload-caps-2026.tabularn<1K0 likes38 downloads17d agoHugging Face26mlfoundations-dev /pdf_qa_r1_annotated_verified_eval_03-18-25_22-14-27_0981tabular1K<n<10K0 likes36 downloads2y agoHugging Face27BEE-spoke-data /napierone-pdf-raw BEE-spoke-data/napierone-pdf-raw NapierOne PDF files converted with marker. detected languages Counter({'en': 4665, 'nl': 2, 'fi': 7, 'fr': 8, 'cy': 54, 'sq': 1, 'it': 1, 'unknown-error': 5, 'sk': 1, 'es': 2, 'de': 3, 'ro': 1, 'pl': 1, 'zh': 1, 'so': 1, 'ml': 1}) tabulartext-generation10K<n<100K0 likes32 downloads9mo agoHugging Face28Baiyinyou /PDF2TEX_raw_withgttabular10K<n<100K0 likes32 downloads10mo agoHugging Face29miracleyin /pdfsys-page-v2-demo pdfsys.page/v2 — 格式演示数据集 pdfsys.page/v2 是 pdfsystem_mnbvc 的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。 这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、 三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。 来源提示:这里的 PDF 页来自 OmniDocBench 与 olmOCR-bench 两个公开 benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些 文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。 一句话设计 一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的; 页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强; 图像像素要么是裁剪图、要么是整页光栅,二选一。 里面有什么 config 行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.tabularimage-to-textn<1K0 likes31 downloads1mo agoHugging Face30Philopater-Luka /Coptic-PDF-Corpustabular1K<n<10K0 likes31 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.