Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes5.4k downloads1y agoHugging Face02sailor2 /sea-pdf-texttext10M<n<100M1 likes1.3k downloads2y agoHugging Face03piushorn /pdf-parse-bench PDF Parse Bench Benchmark for evaluating how effectively PDF parsing solutions extract mathematical formulas and tables from documents. We generate synthetic PDFs with diverse formatting scenarios, parse them with different parsers, and score the extracted content using LLM-as-a-Judge. This semantic evaluation approach substantially outperforms traditional metrics in agreement with human judgment. Leaderboard (2026-Q1) Results are based on two benchmark… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/pdf-parse-bench.documentimage-to-textn<1K2 likes389 downloads7mo agoHugging Face04ankitt6174 /pdf-parse-bench PDF Parse Bench Benchmark for evaluating how effectively PDF parsing solutions extract mathematical formulas and tables from documents. We generate synthetic PDFs with diverse formatting scenarios, parse them with different parsers, and score the extracted content using LLM-as-a-Judge. This semantic evaluation approach substantially outperforms traditional metrics in agreement with human judgment. Leaderboard (2026-Q1) Results are based on two benchmark… See the full description on the dataset page: https://huggingface.co/datasets/ankitt6174/pdf-parse-bench.documentimage-to-textn<1K0 likes387 downloads25d agoHugging Face05Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_pdf llm-jp-corpus-v4 — ja_sip_comprehensive_pdf Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_pdf Files: 156 × jsonl.gz (39.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.texttext-generation1M<n<10M0 likes206 downloads2mo agoHugging Face06Yehor /jfk-pdfs-ocr JFK OCRed files Use colab: https://colab.research.google.com/drive/1xsXHYIaEU50tS0eTwuI5QdkGrbvjHQpW?usp=sharing to get more records. textn<1K0 likes108 downloads2y agoHugging Face07AiAF /SCPWiki-Cleaned-PDF-Archivesstextn<1K0 likes108 downloads1y agoHugging Face08wannaphong /wangchanLION-Curated-pdftext1K<n<10K0 likes97 downloads11mo agoHugging Face09semihk1 /aws-public-pdf-chunked-dataset 📚 AWS PDF Chunk Dataset This dataset consists of chunked text extracted from all publicly available PDF documents on the Amazon Web Services (AWS) official website. The data includes user guides, technical documentation, and best practices about every AWS service, concept, and architecture. It is designed to use in embedding generation, vector databases, and retrieval-augmented generation (RAG) systems. 📦 Dataset Structure The dataset is provided as a .json file.… See the full description on the dataset page: https://huggingface.co/datasets/semihk1/aws-public-pdf-chunked-dataset.text100K<n<1M0 likes68 downloads1y agoHugging Face10NarsAI /pdf-datasettextn<1K0 likes57 downloads3mo agoHugging Face11Baiyinyou /PDF2TEX_raw_withgttabular10K<n<100K0 likes37 downloads10mo agoHugging Face12jellyfish520 /pdf_filestextn<1K0 likes33 downloads3y agoHugging Face13TUaxhuiax /terminal_filesystem_yahoo-finance_pdf-tools_huggingface_1990_settlement_a7c3f91etabularn<1K0 likes32 downloads3d agoHugging Face14Roy229 /pdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_manifest_aurora Aurora Flow - Release Manifest Candidate code name: aurora Version: 2.1.0 Status: stable License: MIT Language: Python Description: Real-time streaming data processing library Maintainer: NovaTech Data Engineering Last release: 2026-07-30 This dataset contains manifest.json (the release manifest) and validate.py (the standardized validation script). Run python validate.py from this directory to validate the release manifest. textn<1K0 likes18 downloads2mo agoHugging Face15Sesamoo /pdf-3dsimulationtextn<1K0 likes17 downloads3y agoHugging Face16Bingdo /design_compiler_pdf_datasettext1K<n<10K0 likes15 downloads2y agoHugging Face17Roy229 /pdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_manifest_crimson Crimson Stream - Release Manifest Candidate code name: crimson Version: 0.9.0 Status: beta License: GPL-3.0 Language: Python Description: Low-latency event streaming engine Maintainer: NovaTech Data Engineering Last release: 2026-08-02 This dataset contains manifest.json (the release manifest) and validate.py (the standardized validation script). Run python validate.py from this directory to validate the release manifest. textn<1K0 likes15 downloads2mo agoHugging Face18Philopater-Luka /Coptic-PDF-Corpustabular1K<n<10K0 likes15 downloads1mo agoHugging Face19cminst /imslp-raw-pdf-collection IMSLP Raw PDF Collection This dataset stores raw public-domain IMSLP PDF downloads collected by ReScore. The raw PDF archive is append-only. Each Modal download launch publishes tar shards under pdf_shards/ and matching JSONL manifests under manifests/. Default collection id: imslp_pdf_collection. The coordinator remains the source of claim/download state. Manifest rows link each tar member back to the IMSLP id, source metadata, PDF checksum, page count, Modal profile, launch id… See the full description on the dataset page: https://huggingface.co/datasets/cminst/imslp-raw-pdf-collection.tabularimage-to-text10K<n<100K0 likes14 downloads2mo agoHugging Face20zhuq41 /filesystem_huggingface_yahoo_pdf_excel_pw_term_schol_3225_pr_a2ttai32textn<1K0 likes14 downloads2mo agoHugging Face21Roy229 /pdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_manifest_bluewave BlueWave ETL - Release Manifest Candidate code name: bluewave Version: 1.4.0 Status: stable License: Apache-2.0 Language: Python Description: Batch ETL orchestration framework Maintainer: NovaTech Data Engineering Last release: 2026-06-15 This dataset contains manifest.json (the release manifest) and validate.py (the standardized validation script). Run python validate.py from this directory to validate the release manifest. textn<1K0 likes14 downloads2mo agoHugging Face22nswamy14 /pdf-frame-dataset-1text10K<n<100K0 likes13 downloads1y agoHugging Face23Sesamoo /for-pdf-jsonltextn<1K0 likes12 downloads3y agoHugging Face24Roy229 /pdf-tools_huggingface_terminal_filesystem_7936_contract_register_t6zf5ptextn<1K0 likes12 downloads2mo agoHugging Face25fbellame /pdf_to_quizz_mistraltextn<1K1 likes11 downloads3y agoHugging Face26Roy229 /pdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_assessment NovaTech Component Assessment Tracker This dataset is the decision tracking repository for the vendor technology assessment. After analysing the proposals and validating each candidate, the analyst must upload the assessment report into this repository as: assessment_report.json assessment_report.csv Acceptance policy: A candidate is APPROVED when its validation passes AND its license is one of MIT or Apache-2.0. Otherwise it is REJECTED. The final recommendation is the… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/pdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_assessment.textn<1K0 likes10 downloads2mo agoHugging Face27alphasoft91 /pdftextn<1K0 likes9 downloads2y agoHugging Face28nlplabtdtu /tdtu_pdf_datatextn<1K0 likes8 downloads3y agoHugging Face29JohnB13 /PDF-FORMStextn<1K0 likes8 downloads3y agoHugging Face30vinodha /pdfjsonltextn<1K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.