Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes6.5k downloads27d agoHugging Face02incrediblecrab /scotus-case-documents Supreme Court of the United States: Case Documents The document files that the Supreme Court's website lists under Case Documents in its footer, one config per collection, each row a file with the file itself, byte for byte, and its text. Nothing here is edited by hand, and no text is corrected, normalized or generated. The pipeline and its tests are in github.com/incrediblecrab/scotus-research-service-products, and this card is rendered from the collections' manifests in the… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/scotus-case-documents.tabular1K<n<10K0 likes3.1k downloads1h agoHugging Face03AiAF /JFK-Assassination-Records-2025-Documents-Release0 likes3.1k downloads2y agoHugging Face04CGIAR /Embrapa-ai-documents-markdown1 likes2.3k downloads3mo agoHugging Face05SherlockRamos /jurisdb-legal-documents JurisDB - Brazilian Legal Documents Dataset Dataset Description This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU). Dataset Structure . ├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/ │ ├── leis_estaduais/ │ ├── leis_federais/ │ └── ... └── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.documenttext-classificationn<1K0 likes2.2k downloads9mo agoHugging Face06Voxel51 /form_understanding_in_noisy_scanned_documents_plus Dataset Card for Form Understanding in Noisy Scanned Documents Plus This is a FiftyOne dataset with 1026 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/form_understanding_in_noisy_scanned_documents_plus") # Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/form_understanding_in_noisy_scanned_documents_plus.imageobject-detection1K<n<10K1 likes1.7k downloads1y agoHugging Face07CGIAR /gardian-cigi-ai-documents-markdown1 likes1.4k downloads3mo agoHugging Face08HumynLabs /French_Documents_Dataset_PDF French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.documentn<1K0 likes1.4k downloads11mo agoHugging Face09th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M46 likes1.3k downloads2mo agoHugging Face10HumynLabs /Arabic_Documents_Dataset_PDF Arabic Documents Dataset (PDF) This dataset contains a collection of Arabic-language documents in PDF format. The corpus includes books, articles, reports, and educational materials written in Modern Standard Arabic and regional variants. It is curated to support AI research in document understanding, Arabic OCR, and text extraction from complex layouts. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Arabic_Documents_Dataset_PDF.documentn<1K0 likes1.2k downloads11mo agoHugging Face11CGIAR /ifpri-ai-documents-markdown GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 21,726 Total Size 623.27 MB Total Tokens 85,359,442 Total Pages 0 Languages 25 Unique Keywords 7,127 Resource Types 20 Date Generated 2026-07-31 02:55:20 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.summarization10K<n<100K1 likes1.1k downloads2mo agoHugging Face12jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes974 downloads1y agoHugging Face13HumynLabs /Japanese_Documents_Dataset_PDF Japanese Documents Dataset (PDF) This dataset contains a curated collection of Japanese-language documents in PDF format. The corpus includes textbooks, research papers, news articles, public-domain books, and government publications written in Japanese. It is intended to support AI research in OCR, document understanding, and multilingual text recognition. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Japanese_Documents_Dataset_PDF.documentn<1K2 likes912 downloads11mo agoHugging Face14kasys /multimodal-tip-of-the-tongue-retrieval-for-scientific-documents Open-Source Scientific Documents This repository contains the scientific-paper corpus and generated artifacts for Multimodal Tip-of-the-Tongue Retrieval for Scientific Papers. It brings together source PDFs, extracted paper content, textual and visual clues, query collections, and split definitions. The generation code documents how these artifacts are made. The PDFs form a shared retrieval corpus. Query and evaluation collections refer to paper identifiers in that corpus; the… See the full description on the dataset page: https://huggingface.co/datasets/kasys/multimodal-tip-of-the-tongue-retrieval-for-scientific-documents.documentvisual-document-retrieval100K<n<1M1 likes874 downloads4d agoHugging Face15kasys /open-source-scientific-documents0 likes798 downloads4d agoHugging Face16ankitjh4 /bharat-government-documents Bharat Guide: screened Indian public information documents This snapshot contains 64,964 distinct normalized text bodies and 77,526 source records. Generated 2026-10-03T01:10:02.179230+00:00. Contents and provenance One Parquet row represents one normalized source body, with original extracted text, source title, URL, extraction quality, observed retrieval time, rule version, topic labels and all current-body source aliases. Duplicate URLs/editions are preserved… See the full description on the dataset page: https://huggingface.co/datasets/ankitjh4/bharat-government-documents.textquestion-answering10K<n<100K59 likes782 downloads6d agoHugging Face17abigailhaddad /sam-solicitation-documents Sam Solicitation Documents Attachments from federal solicitation notices on SAM.gov: statements of work, performance work statements, justifications, amendments, wage determinations and the rest of the paperwork that accompanies a federal contract opportunity. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/sam-solicitation-documents.text-retrieval0 likes736 downloads23d agoHugging Face18HumynLabs /Spanish_Documents_Dataset_PDF Spanish Documents Dataset (PDF) This dataset contains a curated collection of Spanish-language documents in PDF format. It includes books, educational materials, research papers, news articles, and government publications written in Spanish. The dataset supports AI research in OCR, multilingual document understanding, and text extraction for Latin-script languages. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Spanish_Documents_Dataset_PDF.documentn<1K0 likes721 downloads11mo agoHugging Face19HumynLabs /German_Documents_Dataset_PDFdocumentn<1K0 likes708 downloads11mo agoHugging Face20DeyangKong /documents_QA_100docs0 likes705 downloads2y agoHugging Face21singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes694 downloads11mo agoHugging Face22HumynLabs /Italian_Documents_Dataset_PDF Italian Documents Dataset (PDF) This dataset contains a curated collection of Italian-language documents in PDF format. It includes books, academic publications, reports, government documents, and news articles written in Italian. The dataset supports AI research in OCR, multilingual document understanding, and text recognition for Romance languages. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Italian_Documents_Dataset_PDF.documentn<1K0 likes692 downloads11mo agoHugging Face23ZamAI-Pashto /zamai-pashto-documents ZamAI-Pashto Documents This repository contains the dataset and scripts for the ZamAI-Pashto document processing project, focusing exclusively on the Pashto language. Project Structure data/: Contains scanned documents, extracted text, translations, and summaries. annotations/: OCR bounding boxes, handwriting labels, and domain tags. scripts/: OCR processing, text cleaning, and translation alignment scripts. configs/: Dataset configurations and OCR settings.… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-documents.textvisual-document-retrievaln<1K0 likes663 downloads2mo agoHugging Face24hotchpotch /cc100-ja-documents cc100-ja-documents HuggingFace で公開されている cc100 / cc100-ja は line 単位の分割のため、document 単位に結合したものです。 ライセンスはオリジナルのcc100 に準拠します。 text10M<n<100M4 likes634 downloads2y agoHugging Face25nyuuzyou /znanio-documents Dataset Card for Znanio.ru Educational Documents Dataset Summary This dataset contains 588,545 educational document files from the znanio.ru platform, a resource for teachers, educators, students, and parents providing diverse educational content. Znanio.ru has been a pioneer in educational technologies and distance learning in the Russian-speaking internet since 2009. The dataset includes a small portion of English language content, primarily for language learning… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-documents.texttext-classification100K<n<1M0 likes611 downloads2y agoHugging Face26HumynLabs /Chinese_Documents_Dataset_PDF Chinese Documents Dataset (PDF) This dataset consists of a curated collection of Chinese-language documents in PDF format. It includes textbooks, research papers, articles, public-domain books, and official documents written in Simplified and Traditional Chinese. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Chinese_Documents_Dataset_PDF.documentn<1K0 likes585 downloads11mo agoHugging Face27eddmpython /cleangov-local-documents 지방자치단체 예산서·결산서·재정공시·계약현황·세입세출 원문 색인 (지방재정365 "우리 지자체" 링크 API) 과 원문 각 자치단체가 자기 홈페이지에 공개하는 예산서, 결산서(성과보고서·성인지·기금 결산 포함), 재정공시, 계약현황, 세입세출현황 원문의 자치단체별 연도별 URL 색인과 그 원문. 색인은 지방재정365 허브의 "우리 지자체 *" 5 서비스(SETLK 결산서·BUDLK 예산서·FINLK 재정공시·CONLK 계약현황·REVLK 세입세출현황)가 회계연도 × 243 자치단체 × 링크(lnk_url_nm) 로 준다. 원문 자체는 자치단체 서식(PDF·HWP·HWPX·XLSX)이고 세부 산출내역과 판단에 필요한 문맥이 여기 있다. 지방 성과보고서(EVAL-004)와 자체감사·결산검사(AUDIT-003) 도 이 색인으로 각 자치단체의 결산 문서에 도달한다. 출처: https://www.lofin365.go.kr/portal/LF5100000.do 이용… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/cleangov-local-documents.text10K<n<100K0 likes520 downloads6d agoHugging Face28sentence-transformers /example-documents Example Documents A small set of example documents across modalities (image, audio, video) for use in Sentence Transformers retrieval snippets and documentation. These are the kinds of files you pass to model.encode_document(...). They can safely be used as examples in your model cards if you don't want to host the example assets in your model repositories themselves. Contents File Modality doc1.jpg image (document page) doc2.jpg image (document page)… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/example-documents.audion<1K1 likes466 downloads2mo agoHugging Face29KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes411 downloads4mo agoHugging Face30HumynLabs /Russian_Documents_Dataset_PDF Russian Documents Dataset (PDF) This dataset contains a curated collection of Russian-language documents in PDF format. The corpus includes books, academic papers, government publications, articles, and educational materials written in Russian. It is designed to support AI research in OCR, document understanding, and multilingual text recognition. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Russian_Documents_Dataset_PDF.documentn<1K1 likes406 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.