Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AyoubChLin /Company-document-dataset-v2 Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentdocument-question-answering100K<n<1M2 likes12k downloads3d agoHugging Face02mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes3.5k downloads1mo agoHugging Face03incrediblecrab /scotus-case-documents Supreme Court of the United States: Case Documents The document files that the Supreme Court's website lists under Case Documents in its footer, one config per collection, each row a file with the file itself, byte for byte, and its text. Nothing here is edited by hand, and no text is corrected, normalized or generated. The pipeline and its tests are in github.com/incrediblecrab/scotus-research-service-products, and this card is rendered from the collections' manifests in the… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/scotus-case-documents.tabular1K<n<10K0 likes3.1k downloads9h agoHugging Face04RoboCOIN /Agilex_Cobot_Magic_zip_up_the_document_baggated Agilex_Cobot_Magic_zip_up_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: Agilex_Cobot_Magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pull place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.tabularrobotics100K<n<1M0 likes1.5k downloads4mo agoHugging Face05RoboCOIN /RMC-AIDA-L_organise_the_document_baggated RMC-AIDA-L_organise_the_document_bag 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: realman_rmc_aidal | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick pull 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_organise_the_document_bag.tabularrobotics100K<n<1M0 likes1k downloads10mo agoHugging Face06jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes974 downloads1y agoHugging Face07singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes694 downloads11mo agoHugging Face08KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes411 downloads4mo agoHugging Face09dtunkelang /bag-of-documents Bag-of-Documents: Product Search Dataset Blog post: Distilling Retrieval Pipelines to a Single Embedding Model Live demo: huggingface.co/spaces/dtunkelang/bag-of-documents-demo Code: github.com/dtunkelang/bag-of-documents Dataset Description A large-scale bag-of-documents dataset for e-commerce product search, built on Amazon product data. Each search query is represented as a distribution of relevant products in embedding space, captured by a centroid vector and a… See the full description on the dataset page: https://huggingface.co/datasets/dtunkelang/bag-of-documents.tabularsentence-similarityn<1K5 likes405 downloads5mo agoHugging Face10prg-unibe /dodis-historical-documents Dataset Overview The Swiss Historical Archive of DOdis Works (SHADOW) is a large-scale benchmark dataset derived from the resources of Dodis, an independent research center dedicated to the study of Swiss foreign policy and Switzerland’s international relations. Dodis has curated and published nearly 50,000 key documents that document the administrative practices and decision-making processes of the Swiss federal administration. The underlying corpus consists of notes, letters… See the full description on the dataset page: https://huggingface.co/datasets/prg-unibe/dodis-historical-documents.tabularsummarization10K<n<100K5 likes376 downloads4mo agoHugging Face11hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes349 downloads9d agoHugging Face12hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes344 downloads9d agoHugging Face13LegionIntel /named_entity_recognition_document_contexttabular1M<n<10M9 likes335 downloads2y agoHugging Face14DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes264 downloads4mo agoHugging Face15AccountVerify /dutch-legal-documents Dutch Legal Documents A comprehensive collection of 1.14 million Dutch legal documents, including court rulings, parliamentary documents, and EU legislation. Sources Source Type Documents Description Rechtspraak.nl Court rulings 882,212 All Dutch court rulings from Hoge Raad, Raad van State, Gerechtshoven, Rechtbanken, CRvB, CBb Officiële Bekendmakingen Parliamentary docs 255,837 Kamerstukken, Handelingen, Memories van Toelichting, Staatsblad… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/dutch-legal-documents.tabular1M<n<10M1 likes258 downloads21d agoHugging Face16sheggle /dutch-legal-documents Dutch Legal Documents A comprehensive collection of 1.14 million Dutch legal documents, including court rulings, parliamentary documents, and EU legislation. Sources Source Type Documents Description Rechtspraak.nl Court rulings 882,212 All Dutch court rulings from Hoge Raad, Raad van State, Gerechtshoven, Rechtbanken, CRvB, CBb Officiële Bekendmakingen Parliamentary docs 255,837 Kamerstukken, Handelingen, Memories van Toelichting, Staatsblad EUR-Lex EU… See the full description on the dataset page: https://huggingface.co/datasets/sheggle/dutch-legal-documents.tabular1M<n<10M1 likes251 downloads6mo agoHugging Face17erdem-erdem /Turkish-Law-Documents-700k-clustered Turkish Legal Documents Clustering Dataset A comprehensive dataset of 700,000 Turkish legal documents from the two primary sources of legal precedent in Turkey, clustered using multiple emebdding models and algorithms to enable research, analysis, and machine learning applications. Overview This repository contains a large-scale document clustering pipeline and dataset for Turkish legal documents sourced from: Yargıtay - Turkey's highest court of appeal for civil and… See the full description on the dataset page: https://huggingface.co/datasets/erdem-erdem/Turkish-Law-Documents-700k-clustered.tabular100K<n<1M7 likes223 downloads11mo agoHugging Face18trannhiem /TranNhiem-Vietnamese-DocumentImage-Reasoning TranNhiem Vietnamese Document-Image Reasoning (V-Doc) Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5 over the Viet-Doc-VQA-II document collection. Curated by: Trần Nhiệm.. Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.imagevisual-question-answering10K<n<100K5 likes135 downloads3mo agoHugging Face19histde /dta-documents Deutsches Textarchiv (DTA) Documents This datasets hosts all documents from the Deutsches Textarchiv (DTA). One row per work of the Deutsches Textarchiv (DTA), built from the official DTA dump dta_komplett_2026-02-10 (TCF format). The dataset contains 5,480 documents spanning 1472 to 1987 with about 205M tokens and 1.4B characters of German text, with two text views per document: text: the historical, layout-faithful transcription (line breaks, long s ſ, combining diacritics… See the full description on the dataset page: https://huggingface.co/datasets/histde/dta-documents.tabular1K<n<10K0 likes128 downloads2mo agoHugging Face20thoughtworks /document-processing-benchmark Document Processing Benchmark 8 public document datasets (receipts, invoices, forms, bank statements, multi-page docs, contracts) normalized into one parquet schema. Each row has the document, ground-truth annotations, and per-row token/latency/cost numbers from real API calls to one or more reference models. You can read off a target's cost/latency/quality without re-running it. from datasets import load_dataset ds = load_dataset("thoughtworks/document-processing-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/document-processing-benchmark.tabularimage-to-text10K<n<100K1 likes123 downloads5mo agoHugging Face21operant-ai /doclaynet-document-level DocLayNet Document-Level Reconstruction and 8K Expansion This dataset is a normalized, one-row-per-document view over the page-level DocLayNet v1.1 dataset. Pages are grouped using DocLayNet's source metadata and ordered by their original page number. Dataset summary 2,944 logical documents 80,863 observed pages 896 complete document groups 2,048 partial document groups Train: 2,355 documents / 60,810 pages Validation: 294 documents / 7,964 pages Test: 295… See the full description on the dataset page: https://huggingface.co/datasets/operant-ai/doclaynet-document-level.tabular10K<n<100K1 likes110 downloads29d agoHugging Face22ClimatePolicyRadar /global-stocktake-documents Global Stocktake Open Data This repo contains the data for the first UNFCCC Global Stocktake. The data consists of document metadata from sources relevant to the Global Stocktake process, as well as full text parsed from the majority of the documents. The files in this dataset are as follows: metadata.csv: a CSV containing document metadata for each document we have collected. This metadata may not be the same as what's stored in the source databases – we have cleaned and added… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/global-stocktake-documents.tabular1M<n<10M7 likes106 downloads3y agoHugging Face23ClimatePolicyRadar /all-document-text-datagated Climate Policy Radar Open Data This repo contains the full text data of all of the documents from the Climate Policy Radar database (CPR), which is also available at Climate Change Laws of the World (CCLW). Please note that this replaces the Global Stocktake open dataset: that data, including all NDCs and IPCC reports is now a subset of this dataset. What’s in this dataset This dataset contains two corpus types (groups of the same types or sources of documents) which… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/all-document-text-data.tabular10M<n<100M24 likes97 downloads1y agoHugging Face24FudanSELab /SO_KGXQR_DOCUMENT Dataset Card for "SO_KGXQR_DOCUMENT" More Information needed tabular100K<n<1M0 likes94 downloads3y agoHugging Face25evalitahf /document_dating Dating Document Evaluation at EVALITA 2020 In the context of EVALITA 2020, we propose the task of assigning a temporal span to a document, i.e. recognising when a document was issued. The task has already been addressed in other languages, namely French, English, Polish, also in the framework of shared tasks, see for example the DÉfi Fouille de Textes (DEFT) 2010 and 2011 challenges (Grouin, 2010; Grouin, 2011), the SemEval-2015 task on Diachronic Text Evaluation (Popescu and… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/document_dating.tabulartext-classification1K<n<10K0 likes81 downloads2y agoHugging Face26MonumentalSystems /document-corpus-v3-open Document Corpus v3 Open document-corpus-v3-open is the redistribution-compatible slice of the exact byte-level pretraining corpus used by MonumentalSystems' 128M Harmonic GPT experiments. It contains 869,739 filtered documents and 2.192 GB of UTF-8 text before Parquet compression. This is not the complete internal document-corpus-v3. Restricted, unknown-license, and share-alike sources were excluded conservatively. Every included row comes from an upstream dataset whose card… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/document-corpus-v3-open.tabulartext-generation100K<n<1M0 likes79 downloads2mo agoHugging Face27datamatters24 /research-document-archive Research Document Archive Computational analysis of 234,630 declassified U.S. government documents across 7 archival collections. Output of a 13-step ML pipeline extracting OCR text, entities, topics, keywords, redactions, and semantic embeddings from 3.1 million pages. Files File Rows Description documents.parquet 234,630 Document metadata: id, source_section, file_path, file_hash, total_pages pages/<section>.parquet 3.1M Per-page OCR text + 1536-dim… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/research-document-archive.tabulartext-classification10M<n<100M0 likes69 downloads6mo agoHugging Face28juliensimon /agent-traces-legal-document-analysis Agent Traces: legal-document-analysis Synthetic multi-agent workflow traces with LLM-enriched content for the legal-document-analysis domain. Part of the juliensimon/open-agent-traces collection — 10 datasets covering diverse domains and workflow patterns. What is this dataset? This dataset contains 1,498 events across 50 workflow runs, each representing a complete multi-agent execution trace. Every trace includes: Agent reasoning — chain-of-thought for each agent step… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/agent-traces-legal-document-analysis.tabular1K<n<10K1 likes61 downloads7mo agoHugging Face29JBrightmanAI /TranNhiem-Vietnamese-DocumentImage-Reasoning TranNhiem Vietnamese Document-Image Reasoning (V-Doc) Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B over the Viet-Doc-VQA-II document collection. Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.imagevisual-question-answering10K<n<100K0 likes61 downloads3mo agoHugging Face30QIRIM /crh-parallel-corpora-document-level-noisytabulartranslation10K<n<100K1 likes52 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.