forms
Datasets
All datasets matching “forms”FormStruct-Bench
FormStruct-Bench
Dataset Description
FormStruct-Bench is a multilingual benchmark for extracting the semantic and
spatial structure of forms from document images. The repository combines a
7,000-page main benchmark, a controlled visual-degradation set, and
template-level layout annotations. It supports evaluation of vision-language
models and document AI systems on hierarchical key-value extraction, document
structure recovery, region localization, table and… See the full description on the dataset page: https://huggingface.co/datasets/D2I-CUHK-Shenzhen/FormStruct-Bench.Intel-WebCorpus-forms
💻 Intel WebCorpus Forms (Enterprise Hardware Q&A)
This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums.
It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.cua-s1-forms
cua-s1-forms (dataset)
Synthetic + real training/eval data for cua-ai/cua-s1-forms,
a jev-like one-pass option scorer for GUI form filling behind
cua-driver.
Generator source: cua_s1/synth.py in
https://github.com/trycua/cua/tree/main/libs/cua-s1.
Files
file
rows
source
train.jsonl
~150k
synthetic
validation.jsonl
~18k
synthetic
test.jsonl
~20k
synthetic, form-signature-disjoint from train/validation
demo.jsonl
196
real: 3 real JevBrowser form… See the full description on the dataset page: https://huggingface.co/datasets/cua-ai/cua-s1-forms.irs-formsindian_dance_formsThis dataset is taken from https://www.kaggle.com/datasets/aditya48/indian-dance-form-classification but is originally from the Hackerearth deep learning contest of identifying Indian dance forms. All the credits of dataset goes to them.
Content
The dataset consists of 599 images belonging to 8 categories, namely manipuri, bharatanatyam, odissi, kathakali, kathak, sattriya, kuchipudi, and mohiniyattam. The original dataset was quite unstructured and all the images were put together.… See the full description on the dataset page: https://huggingface.co/datasets/tanmaykm/indian_dance_forms.ANC-forms
