Information Extraction
rst-information-extraction-11bbpmn-information-extraction-v2turkextract-turkish-information-extractioninformation-extraction-llama3-8B-4bit-finetuned-mergednamed-entity-recognition-for-information-extractioninformation-extraction-llama3-8B-qlora128-finetuned-mergedgpt2-delivery-information-extractiondonut-base-Medical_Handwritten_Prescriptions_Information_Extraction_1
key_information_extractionposter-schedule-information-extraction
Multimodal Visual-Text Dataset for Poster Schedule Information Extraction
A ready-to-train Indonesian document AI dataset combining pixels, OCR tokens, spatial layout, and BIO entity labels.
Overview
What it is
127 Indonesian seminar and religious-study event posters with multimodal token-level annotations
Primary task
Schedule information extraction as token classification
Modalities
Image + text + 2D spatial layout
Coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Ibnuck/poster-schedule-information-extraction.kleister_nda_information_extraction
Kleister NDA — Information Extraction (orgrctera/kleister_nda_information_extraction)
Overview
This release packages the Kleister NDA split of the Kleister benchmark as rows suitable for information extraction (IE) evaluation. Each example points at a Non-Disclosure Agreement (NDA) document and specifies which attribute keys should be filled; the target is a JSON object of normalized string values for those keys (with null when a value is absent or not applicable).… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/kleister_nda_information_extraction.html_to_json_information_extraction_dataset
HTML to JSON Information Extraction Dataset
Description
The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON.
These HTML have been sourced (scraped) from about 25 companies' career pages.
The dataset contains three splits - train, test, unseen_test.
This dataset has been built to fine tune SLMs & LLMs for the information extraction task.
train split
This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.Work-At-Height-Safety-Observation-Hazard-Information-Extraction-Dataset
Work-at-Height Safety Observation Hazard Information Extraction Dataset
This dataset contains safety inspection and hazard observation records from work-at-height settings, paired with JSON annotations for work activities, hazard sources, exposed persons or entities, locations, and control measures explicitly mentioned in the records. The paired examples map field observations to structured information and can support training and evaluation of models that interpret safety… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Work-At-Height-Safety-Observation-Hazard-Information-Extraction-Dataset.japanese-confidential-information-extraction-sft
Japanese Confidential Information Extraction — SFT Dataset
日本語テキストから社外秘の固有表現を抽出するタスク向けの SFT (Supervised Fine-Tuning) データセットです。
LFM2 系モデルの LoRA fine-tune を想定して構築されています。
タスク概要
入力テキスト(日本語)に含まれる機密情報を、11カテゴリの JSON として抽出します。
入力: 「山田太郎(yamada@example.co.jp)から請求書番号 INV-2024-0042 で
売上 ¥12,800,000 の見積書が届いた。」
出力: {
"address": [],
"company_name": [],
"email_address": ["yamada@example.co.jp"],
"human_name": ["山田太郎"],
"phone_number": [],
"account_identifier":… See the full description on the dataset page: https://huggingface.co/datasets/akiFQC/japanese-confidential-information-extraction-sft.
