Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01D2I-CUHK-Shenzhen /FormStruct-Bench FormStruct-Bench Dataset Description FormStruct-Bench is a multilingual benchmark for extracting the semantic and spatial structure of forms from document images. The repository combines a 7,000-page main benchmark, a controlled visual-degradation set, and template-level layout annotations. It supports evaluation of vision-language models and document AI systems on hierarchical key-value extraction, document structure recovery, region localization, table and… See the full description on the dataset page: https://huggingface.co/datasets/D2I-CUHK-Shenzhen/FormStruct-Bench.imageimage-to-text1K<n<10K1 likes14k downloads5d agoHugging Face02sphita /Intel-WebCorpus-forms 💻 Intel WebCorpus Forms (Enterprise Hardware Q&A) This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums. It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.textquestion-answering100K<n<1M4 likes830 downloads12d agoHugging Face03cua-ai /cua-s1-forms cua-s1-forms (dataset) Synthetic + real training/eval data for cua-ai/cua-s1-forms, a jev-like one-pass option scorer for GUI form filling behind cua-driver. Generator source: cua_s1/synth.py in https://github.com/trycua/cua/tree/main/libs/cua-s1. Files file rows source train.jsonl ~150k synthetic validation.jsonl ~18k synthetic test.jsonl ~20k synthetic, form-signature-disjoint from train/validation demo.jsonl 196 real: 3 real JevBrowser form… See the full description on the dataset page: https://huggingface.co/datasets/cua-ai/cua-s1-forms.texttext-classification100K<n<1M16 likes792 downloads22d agoHugging Face04TrevorJS /irs-formstext1K<n<10K5 likes420 downloads2y agoHugging Face05precisit /one-pass-sv-forms-synthetic One-Pass SV-Forms synthetic corpus (Swedish) This dataset is entirely synthetic. It was generated by a script from a concept catalogue, not collected from people or customer submissions. Names and organisations are constructed, email addresses use .invalid, and identifier-shaped values are generated locally. They are not checked against registries; coincidental matches with real names or identifiers cannot be ruled out. The dataset is published so the recipe behind… See the full description on the dataset page: https://huggingface.co/datasets/precisit/one-pass-sv-forms-synthetic.texttext-classification100K<n<1M0 likes105 downloads19d agoHugging Face06Victorgonl /UFLA-FORMS UFLA-FORMS: an Academic Forms Dataset for Information Extraction in the Portuguese Language About UFLA-FORMS is a manually labeled dataset of document forms in Brazilian Portuguese extracted from the domains of the Federal University of Lavras (UFLA). The dataset emphasizes the hierarchical structure between the entities of a document through their relationships, in addition to the extraction of key-value pairs. Samples were labeled using ToolRI. Overview… See the full description on the dataset page: https://huggingface.co/datasets/Victorgonl/UFLA-FORMS.imagetoken-classificationn<1K1 likes92 downloads2y agoHugging Face07ift /handwriting_forms Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ift/handwriting_forms.imagefeature-extraction1K<n<10K14 likes80 downloads3y agoHugging Face08hyturing /US_tax_forms_donut NIST-SD2 (US Tax Forms) Donut Dataset Processed version of NIST Special Database 2 for document understanding tasks, formatted for use with the Donut architecture. Contains 5,590 annotated document images (5,031 train, 559 test) across 20 tax form classes. Description Curated by: National Institute of Standards and Technology (NIST) License: MIT Total Size: 947.33 MB Annotations: Class labels (20 tax form types) Full text ground truth image1K<n<10K4 likes64 downloads2y agoHugging Face09Symage /synthetic-us-forms-preview SymageDocs — Synthetic US Forms Preview A small, CC-BY-4.0, fully synthetic document-AI training set: 525 labeled page images across six families of US business and government forms, each page shipping FUNSD ground truth plus a LayoutLM-ready token/bbox/tag view. This is a preview subset. It exists so you can load real output from the SymageDocs generator, inspect the label quality, and decide whether generating your own corpus is worth your time — without an account, an email… See the full description on the dataset page: https://huggingface.co/datasets/Symage/synthetic-us-forms-preview.imagetoken-classificationn<1K2 likes63 downloads2mo agoHugging Face10rumike7 /cadquery-creating-basic-2d-and-3d-formstabular10K<n<100K2 likes59 downloads1y agoHugging Face11forms-legal /legal-templates-multilingual Forms Legal — Multilingual Legal Templates Corpus 19,419 legal document templates across 36 jurisdictions in 12 languages, expanded to 22,876 (document × locale) rows. Released under CC-BY-4.0 by forms-legal.com. Quick description A multilingual corpus of structured legal document templates spanning 36 jurisdictions. Each document includes a multi-section editorial brief (whatIs, whenNeeded, keyElements, howToFill, legalRequirements, commonMistakes), 5-8… See the full description on the dataset page: https://huggingface.co/datasets/forms-legal/legal-templates-multilingual.texttext-generation10K<n<100K0 likes43 downloads4mo agoHugging Face12saurabh1896 /OMR-forms Dataset Card for "OMR-forms" More Information needed imagen<1K1 likes23 downloads3y agoHugging Face13dicta-il /hebrew_suffix_verbal_forms Suffixed Verbal Forms Detection Dataset for Modern Hebrew Dataset Summary This dataset contains annotated Hebrew sentences containing verbal forms that are ambiguous as to whether they include a pronominal suffix or not (e.g., the Hebrew word lamed-yod-mem-daled-vav can be understood as either "he taught him" or "they taught"). The goal of the dataset is to support tasks involving the identification and disambiguation of verbs with pronominal suffixes in Hebrew literature… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/hebrew_suffix_verbal_forms.tabular1K<n<10K0 likes23 downloads2y agoHugging Face14monodox /theyyam-forms-kbtextn<1K0 likes20 downloads6mo agoHugging Face15Symage /coherent-forms-1040-cms1500-i9gated SymageDocs — Coherent US Tax / Health / Employment Forms (FUNSD) A fully synthetic document-AI training set: three US forms — IRS Form 1040, CMS-1500, and USCIS Form I-9 — filled from the same synthetic identity, so name / SSN / address / employer flow consistently across all three renderings. Each page ships with FUNSD ground truth (word boxes, entity labels, key–value linking) plus a LayoutLMv3-ready token/bbox/tag view. 3,000 page-level image + annotation rows (train 2,400 /… See the full description on the dataset page: https://huggingface.co/datasets/Symage/coherent-forms-1040-cms1500-i9.imagetoken-classification1K<n<10K4 likes20 downloads3mo agoHugging Face16Voidreaper2026 /uk-benefit-forms-structured UK Benefit Forms Structured Dataset A structured dataset of 120 UK government benefit and legal forms, extracted and processed for use in AI-assisted form-filling applications. Built as part of the EasyClaimAI project. Why This Dataset Exists Millions of people in the UK struggle with complex government forms — benefit claims, legal applications, pension forms. The language is dense, the guidance is buried, and mistakes can cost people money or delay vital support. This… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/uk-benefit-forms-structured.textquestion-answering1K<n<10K1 likes18 downloads5mo agoHugging Face17Elliot-Data /handwriting_forms_cleanedgated handwriting_forms_cleaned The handwriting_forms__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 1,360 QA turns 5,198 answers rewritten by the cleaning pass 166 QA created by the cleaning pass (new_qa) 3,860 (74.3%) shards 1 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/handwriting_forms_cleaned.imagevisual-question-answering1K<n<10K0 likes17 downloads1mo agoHugging Face18hieunguyen1053 /qa-formstext1K<n<10K0 likes15 downloads2y agoHugging Face19Ronysalem /medical-forms-datasetimagen<1K0 likes15 downloads2y agoHugging Face20electricsheepafrica /africa-uganda-experience-of-various-forms-of-violence-801c079c Experience of Various Forms of Violence | Africa (Uganda Bureau of Statistics) 7 rows - 1 Africa country/area - 2025 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 7 rows from Uganda Bureau of Statistics, covering Experience of Various Forms of Violence. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-experience-of-various-forms-of-violence-801c079c.tabulartabular-classificationn<1K0 likes14 downloads2mo agoHugging Face21awacke1 /LOINC-Panels-and-Formstext1K<n<10K0 likes13 downloads4y agoHugging Face22Process-Venue /Sanskrit-verb-forms Sanskrit Verb Forms Dataset (संस्कृत धातु रूप संग्रह) Overview description: | A comprehensive dataset containing Sanskrit verb conjugations (dhatu roop) with 10,348 unique entries. Each entry provides the complete information about a Sanskrit verb form, including: धातु (Dhatu): The root verb पद (Pada): Voice of the verb (परस्मैपद/आत्मनेपद) लकार (Lakara): Tense/mood of the verb पुरुष (Purusha): Person (प्रथम/मध्यम/उत्तम) वचन (Vachana): Number (एकवचन/द्विवचन/बहुवचन)… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-verb-forms.texttext-classification10K<n<100K1 likes12 downloads1y agoHugging Face23Sukuna404 /arabic-legal-documents-and-verified-formstext1K<n<10K0 likes12 downloads4mo agoHugging Face24electricsheepafrica /africa-uganda-forms-of-spousal-violence-4ca2303e Forms of Spousal Violence | Africa (Uganda Bureau of Statistics) 8 rows - 1 Africa country/area - 2025 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 8 rows from Uganda Bureau of Statistics, covering Forms of Spousal Violence. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures Official statistics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-forms-of-spousal-violence-4ca2303e.tabulartabular-classificationn<1K0 likes12 downloads2mo agoHugging Face25elliot-mllm /handwriting_forms_cleanedgated handwriting_forms_cleaned The handwriting_forms__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 1,360 QA turns 5,198 answers rewritten by the cleaning pass 166 QA created by the cleaning pass (new_qa) 3,860 (74.3%) shards 1 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/handwriting_forms_cleaned.imagevisual-question-answering1K<n<10K0 likes11 downloads1mo agoHugging Face26nnul /forms-from-rvl-cdipimage10K<n<100K0 likes9 downloads1y agoHugging Face27JohnB13 /PDF-FORMStextn<1K0 likes8 downloads3y agoHugging Face28avfattakhova /fixed_forms Dataset Card for "fixed_forms" We propose a new dataset pf fixed poetic forms in English, which can be used both for literary analysis and for training of poetry generators (including large language models). The general structure of the dataset is as follows: it contains 12 rows - according to the number of fixed forms, which are: ballade rondeau triolet ottava rima italian (petrarchan) sonnet french sonnet english (shakespearean) sonnet ode stanza elegiac distich (couplet) haiku… See the full description on the dataset page: https://huggingface.co/datasets/avfattakhova/fixed_forms.textn<1K0 likes8 downloads1y agoHugging Face29electricsheepafrica /africa-ilo-pop-3wap-sex-age-tra-nb-youth-working-age-population-by-sex-age-and-forms Youth working-age population by sex, age and forms of transition (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-pop-3wap-sex-age-tra-nb-youth-working-age-population-by-sex-age-and-forms.tabulartabular-classification10K<n<100K0 likes8 downloads2mo agoHugging Face30Ronysalem /donut-medical-formstextn<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.