Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AyoubChLin /Company-document-dataset-v2 Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentdocument-question-answering100K<n<1M2 likes12k downloads3d agoHugging Face02th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M46 likes1.3k downloads2mo agoHugging Face03singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes694 downloads11mo agoHugging Face04Lukaszl /clearocr-invoice-document-ai clearOCR Invoice Document AI Dataset This dataset shows a complete invoice document AI workflow built around clearOCR. It contains 423 high-confidence invoice examples with: original invoice images, OCR text generated by clearOCR, Markdown reconstruction of the document, structured invoice JSON generated by a local fine-tuned extraction model, visual verification metadata. The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.imageimage-to-textn<1K0 likes509 downloads5mo agoHugging Face05KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes411 downloads4mo agoHugging Face06hulk10 /conseil-administratives-appel-full-documents Décisions de Justice Administrative Françaises Description du dataset Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.texttext-classification1K<n<10K1 likes399 downloads9d agoHugging Face07hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes344 downloads9d agoHugging Face08ISLAM-PO /documents-Egyptian-Arabic Dataset evaluation: See EVALUATION.md for schema/config checks, row-count status, and the data-governance plan. Viewer note: default is a lightweight preview; select egyptian_arabic or another source config to load the full data. Current Hub Validation Status Dataset Server rows: 25,399,945 Dataset Server original/Parquet size: 2,758,228,707 bytes (~2.76 GB) Default Hub configuration currently exposes one column: text Source-specific configurations are explicitly declared in… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.texttranslationn<1K2 likes308 downloads15d agoHugging Face09DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes264 downloads4mo agoHugging Face10mtybilly /apex-r1-real-world-documents Apex-R1 Real-World Benchmark Documents This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation. The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks. Contents benchmark_documents/ EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.documentdocument-question-answeringn<1K0 likes256 downloads3mo agoHugging Face11vohuutridung /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.texttext-classification1M<n<10M3 likes241 downloads7mo agoHugging Face12YuITC /Vietnamese-Legal-Documents Vietnamese Legal Documents Dataset 1. Dataset Summary Raw data: tmnam20/BKAI-Legal-Retrieval The Vietnamese Legal Documents Dataset is a benchmark dataset designed for legal information retrieval in the Vietnamese language. It consists of: A corpus of legal documents. Train/test splits containing natural language queries and their corresponding relevant documents. This dataset is intended to support research and development in: Information Retrieval (IR)… See the full description on the dataset page: https://huggingface.co/datasets/YuITC/Vietnamese-Legal-Documents.texttext-retrieval100K<n<1M6 likes207 downloads7mo agoHugging Face13jeffmeloy /python_documentation_codeDataset created using https://github.com/jeffmeloy/py2dataset using the Python code from the following: https://github.com/ansible/ansible https://github.com/apache/airflow https://github.com/arogozhnikov/einops https://github.com/arviz-devs/arviz https://github.com/astropy/astropy https://github.com/biopython/biopython https://github.com/bjodah/chempy https://github.com/bokeh/bokehhttps://github.com/CalebBell/thermo https://github.com/camDavidsonPilon/lifelines https://github.com/coin-or/pulp… See the full description on the dataset page: https://huggingface.co/datasets/jeffmeloy/python_documentation_code.texttext-generation10K<n<100K1 likes156 downloads2y agoHugging Face14minhnguyent546 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/vietnamese-legal-documents.texttext-classification1M<n<10M2 likes139 downloads7mo agoHugging Face15lukesjordan /worldbank-project-documents Dataset Card for World Bank Project Documents Dataset Summary This is a dataset of documents related to World Bank development projects in the period 1947-2020. The dataset includes the documents used to propose or describe projects when they are launched, and those in the review. The documents are indexed by the World Bank project ID, which can be used to obtain features from multiple publicly available tabular datasets. Supported Tasks and Leaderboards No… See the full description on the dataset page: https://huggingface.co/datasets/lukesjordan/worldbank-project-documents.texttable-to-text10K<n<100K5 likes132 downloads4y agoHugging Face16PersianML /persian-document-corpus Persian Document Corpus Dataset Summary The Persian Document Corpus (PDC) is a large collection of Persian documents, comprising over 13,000 files, gathered from publicly accessible PDFs across a wide array of knowledge domains. This corpus includes research articles, theses, dissertations, scientific reports, and book chapters, offering a rich and diverse resource for the Persian Natural Language Processing (NLP) community. It is designed to facilitate the… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-document-corpus.texttext-generation10K<n<100K0 likes109 downloads3mo agoHugging Face17OO-LD /oold-wikidata-schemaorg-documents Wikipedia leads, as the corpus measured them The article leads that oold-llm-bench's Wikidata-schema.org corpus cites: 959 documents, one per entity, each the plain-text introduction of a named revision with its whitespace collapsed. Published because the alternative does not work. The corpus commits each lead's revision id and sha256 rather than its text, and a third party was meant to re-fetch. Extracts are not versioned: the MediaWiki API ignores revids for prop=extracts and… See the full description on the dataset page: https://huggingface.co/datasets/OO-LD/oold-wikidata-schemaorg-documents.texttext-generationn<1K0 likes93 downloads4d agoHugging Face18MonumentalSystems /document-corpus-v3-open Document Corpus v3 Open document-corpus-v3-open is the redistribution-compatible slice of the exact byte-level pretraining corpus used by MonumentalSystems' 128M Harmonic GPT experiments. It contains 869,739 filtered documents and 2.192 GB of UTF-8 text before Parquet compression. This is not the complete internal document-corpus-v3. Restricted, unknown-license, and share-alike sources were excluded conservatively. Every included row comes from an upstream dataset whose card… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/document-corpus-v3-open.tabulartext-generation100K<n<1M0 likes79 downloads2mo agoHugging Face19Shekswess /legal-documents Description Topic: Legal Documents Domains: Law, Contracts, Regulations Focus: Synthetic raw text data for legal document analysis Number of Entries: 1000 Dataset Type: Raw Dataset Model Used: bedrock/us.amazon.nova-pro-v1:0 Language: English Generated by: SynthGenAI Package texttext-generation1K<n<10K8 likes77 downloads1y agoHugging Face20pdt590 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 1,335 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/pdt590/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes67 downloads7mo agoHugging Face21anhquan12 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn Language:… See the full description on the dataset page: https://huggingface.co/datasets/anhquan12/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes67 downloads6mo agoHugging Face22Kemsekov /Corrupted-russian-word-documents-text-datasetThis is synthetic text-corruption dataset, based on Russian open official/government documents text. This dataset is intended to be used to train LLM to perform text-recovery task. All the errors in text is made solely in Russian sentences, hence ignoring any English sentence. Texts contains complex formatting, which is common for documents. Each line contains json object that have array messages value, which consists of role-based conversation. Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.texttext-generationn<1K2 likes52 downloads2y agoHugging Face23lovienal /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 1,335 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/lovienal/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes48 downloads7mo agoHugging Face24timodonnell /bioreason-pro-sft-reasoning-documents BioReason-Pro SFT Reasoning Documents Complete, self-contained training documents reconstructed from the BioReason-Pro SFT data, ready for LLM pre-training. The upstream dataset wanglab/bioreason-pro-sft-reasoning-data ships the assistant side of each training example (reasoning, final_answer) alongside the raw biological context columns, but not the assembled prompt. The prompt cannot be recovered from the data card alone, because two of its three parts were non-textual… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/bioreason-pro-sft-reasoning-documents.tabulartext-generation100K<n<1M0 likes46 downloads2mo agoHugging Face25beatsprom /multimodal-vision-ocr-document-parsing-2026 📐 Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026) This repository provides the official 100-sample production teaser of the Multimodal Vision-Language & Industrial OCR Document Parsing SFT/DPO Suite (2026) by BeatsProm AI Research Lab. The dataset is engineered to train open-weights Vision-Language Models (Qwen2-VL, Pixtral-12B, Llama-3.2-Vision, ColPali) on dense document parsing, normalized spatial bounding boxes (<box>[ymin, xmin, ymax… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-ocr-document-parsing-2026.textimage-to-textn<1K0 likes46 downloads1mo agoHugging Face26Teen-Different /grpo-oumi-synthetic-document-claims Dataset Card for GRPO Oumi ANLI Subset Dataset This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer. You can find more detailed information about the original dataset at the provided link. Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims Dataset Structure The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.texttext-generation1K<n<10K0 likes43 downloads1y agoHugging Face27vanhthefirst /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn Language:… See the full description on the dataset page: https://huggingface.co/datasets/vanhthefirst/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes42 downloads6mo agoHugging Face28mshojaei77 /persian-document-corpus Persian Document Corpus Dataset Summary The Persian Document Corpus (PDC) is a large collection of Persian documents, comprising over 13,000 files, gathered from publicly accessible PDFs across a wide array of knowledge domains. This corpus includes research articles, theses, dissertations, scientific reports, and book chapters, offering a rich and diverse resource for the Persian Natural Language Processing (NLP) community. It is designed to facilitate the training and… See the full description on the dataset page: https://huggingface.co/datasets/mshojaei77/persian-document-corpus.texttext-generation10K<n<100K1 likes39 downloads2y agoHugging Face29stindardlogic /document-summarization-dpo-100k Document Summarization DPO (100K) 100,000 DPO (Direct Preference Optimization) preference pairs for training models to summarize business and professional documents with precision, structure, and analytical depth. Motivation Document summarization is one of the highest-value enterprise AI applications — analysts, lawyers, product managers, and executives use AI to process reports, contracts, and research daily. Models commonly fail by: Losing quantitative data:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/document-summarization-dpo-100k.texttext-generation100K<n<1M0 likes39 downloads3mo agoHugging Face30playforgecoding /spelling-creator-document-import Spelling Creator document import One section of a lesson as someone might have typed it up, paired with that section as lesson JSON. Spelling Creator is a lesson builder for Spelling to Communicate (S2C), used with nonspeaking spellers. Its Import from text reads a typed lesson with rules, and hands the sections the rules cannot read to a small on-device model; this dataset trains that model. What is in it Chat-format JSONL, ready for a supervised fine-tune (TRL's… See the full description on the dataset page: https://huggingface.co/datasets/playforgecoding/spelling-creator-document-import.texttext-generationn<1K0 likes35 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.