datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
skill-extraction-tech
Skill Extraction with ESCO skills - TECH subset
Dataset Summary
This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-tech.skill-extraction-house
Skill Extraction with ESCO skills - HOUSE subset
Dataset Summary
This dataset contains an extension of the HOUSE subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-house.skill-extraction-techwolf
Skill Extraction with ESCO skills - TechWolf subset
Dataset Summary
The TECHWOLF subset, although smaller, represents a more generic distribution of job descriptions and skill spans. ESCO skills are directly annotated on the full sentence level, thus omitting the intermediate span identification step. ESCO v1.1.0 is used.
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-techwolf.insurance-claims-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here:
https://github.com/cleanlab/structured-output-benchmark/
fire-financial-ner-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here:
https://github.com/cleanlab/structured-output-benchmark/
keyphrase_extraction_russianA dataset for extracting or generating keywords for scientific texts in Russian. It contains annotations of scientific articles from four domains: mathematics and computer science (a fragment of the Keyphrases CS&Math Russian dataset), history, linguistics, and medicine. For more details about the dataset please refer the original paper: https://arxiv.org/abs/2409.10640
Dataset Structure
abstract, an abstract in a string format;
keywords, a list of keyphrases provided by… See the full description on the dataset page: https://huggingface.co/datasets/aglazkova/keyphrase_extraction_russian.pii-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here:
https://github.com/cleanlab/structured-output-benchmark/
autonomous-driving-intention-field-extraction-v0.1What this dataset tests
Whether a system can infer agent intentions
from context cues in complex driving scenes.
This is not trajectory prediction.
It is intention inference.
Required outputs
agent_id
inferred_intention
intention_confidence
time_horizon_s
alternative_intentions
stability_score
Scoring conventions
confidence and stability range 0 to 1
time horizon is seconds into the near future
Use case
Layer one of Intention Field and Social Coherence Maps.
This enables… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-intention-field-extraction-v0.1.Table-Extraction
Table Extract Dataset
This dataset is designed to evaluate the ability of large language models (LLMs) to extract tables from text. It provides a collection of text snippets containing tables and their corresponding structured representations in JSON format.
Source
The dataset is based on the Table Fact Dataset, also known as TabFact, which contains 16,573 tables extracted from Wikipedia.
Schema:
Each data point in the dataset consists of two elements:… See the full description on the dataset page: https://huggingface.co/datasets/Effyis/Table-Extraction.joa-extraction-arena
JOA Job-Posting Extraction Arena
Disclosure: Job Opportunities API (JOA, jobopportunitiesapi.org) is an independent data business that sells API access to job-posting data. AI helped run the experiments, check the numbers and draft this text; Loukas (Luca) Tzekos is editorially responsible. Contact: hello@jobopportunitiesapi.org.
A benchmark for one narrow task: reading a real job posting and filling 11 structured fields. Job Opportunities API (JOA) built it to choose and train… See the full description on the dataset page: https://huggingface.co/datasets/JobOpportunitiesAPI/joa-extraction-arena.clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1What this dataset tests
Whether a model can extract the minimal cross-system coherence factor setthat explains resilience or vulnerability.
It rewards
minimal factor selection
correct coupling recognition
ranking by dominance
Coherence factor labels
buffering_capacity_high
buffering_capacity_low
variance_damping_high
variance_damping_low
autonomic_inflammatory_coupling
sleep_metabolic_coupling
stress_inflammation_coupling
immune_metabolic_instability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1.Multilingual_Topic-Specific_Article-Extraction_and_Classification
Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset
This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text.
Cite the Dataset
Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.Text-Classification-and-Relation-Event-Extraction-Mix-datasetsThe paper of GIELLM dataset.
https://arxiv.org/abs/2311.06838
Cite:
@article{gan2023giellm,
title={Giellm: Japanese general information extraction large language model utilizing mutual reinforcement effect},
author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori},
journal={arXiv preprint arXiv:2311.06838},
year={2023}
}
The dataset constructed base in livedoor news corpus 関口宏司 https://www.rondhuit.com/download.html
customer_service_information_extractionclinical-systemic-drift-fingerprint-extraction-v0.1What this dataset tests
The whole-system drift patterninduced by a single drugover time.
Required outputs
drift vector
affected system axes
temporal profile
reversibility
net coherence change
Use case
Foundation layer for the Drug-Induced System Drift Library.
skill-extraction-tech
Skill Extraction with ESCO skills - TECH subset
Dataset Summary
This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/Balki16/skill-extraction-tech.vie-news-tags-extractionpage_extraction_dataset
Title
Page Extraction Dataset
Description
In digitised cultural heritage items such as books, newspapers and archival records, a problem that can negatively affect OCR are black margins around a page caused by document scanning. In order to enable document layout analysis (DLA), these black margins need to be cropped and the pages need to be extracted correctly. To enable the training of a machine learning model capable of extracting pages, a dataset was created. The… See the full description on the dataset page: https://huggingface.co/datasets/SBB/page_extraction_dataset.FEVER_claim_extractionI found this dataset on my harddrive, which if I remember correctly I got from the source mentioned in the paper:
"Claim extraction from text using transfer learning" - By Acharya Ashish Prabhakar, Salar Mohtaj, Sebastian Möller
https://aclanthology.org/2020.icon-main.39/
The github repo with the data seems down.
It extends FEVER dataset with non-claims for training claim detectors.
Drug_Combination_ExtractionBiomedical_Entity_Relation_Extractionclinical-crisis-invariant-signature-minimal-set-extraction-v0.1What this dataset tests
Whether a model can extract the smallest cross-scale signature setthat marks the onset of a clinical crisis.
It penalizes long feature lists.
Outputs
signature_set_top3
dominance_rank_order
mechanism_hypothesis_1_sentence
Signature component labels
coupling_direction_flip
variance_jump_no_threshold
medication_response_mismatch
narrative_alarm_language_shift
lab_lag_reversal_pattern
vitals_labs_decoupling
stealth_hypoperfusion_pattern… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-crisis-invariant-signature-minimal-set-extraction-v0.1.semantic_relations_extraction
Dataset Card for "Semantic Relations Extraction"
Dataset Description
Repository
The "Semantic Relations Extraction" dataset is hosted on the Hugging Face platform, and was created with code from this GitHub repository.
Purpose
The "Semantic Relations Extraction" dataset was created for the purpose of fine-tuning smaller LLama2 (7B) models to speed up and reduce the costs of extracting semantic relations between entities in texts. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/DehydratedWater42/semantic_relations_extraction.digit_extraction
Dataset Card for Dataset Name
Dataset Details
Dataset Description
This dataset is designed to extract and record individual digits from numbers. For each number, it contains descriptions of the digit being extracted, such as the 1st digit, 2nd digit, etc., and the corresponding extracted digit. This dataset can be used for tasks such as training digit extraction models, digit recognition, and sequence-based learning tasks.
Curated by: BEN SLAMA Farah… See the full description on the dataset page: https://huggingface.co/datasets/farahbs/digit_extraction.event_extractionclinical-entity-extraction-transformedmain_product_extractiont5-metadata-extraction-dataset1t5-metadata-extraction-dataset2Triplets-Extraction2
