Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TechWolf /skill-extraction-tech Skill Extraction with ESCO skills - TECH subset Dataset Summary This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0). This dataset is part of a three-part evaluation dataset for skill extraction: skill-extraction-tech skill-extraction-house skill-extraction-techwolf Citation Information If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-tech.texttext-classification1K<n<10K2 likes733 downloads2y agoHugging Face02TechWolf /skill-extraction-house Skill Extraction with ESCO skills - HOUSE subset Dataset Summary This dataset contains an extension of the HOUSE subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0). This dataset is part of a three-part evaluation dataset for skill extraction: skill-extraction-tech skill-extraction-house skill-extraction-techwolf Citation Information If you use this dataset, please… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-house.texttext-classification1K<n<10K5 likes576 downloads2y agoHugging Face03TechWolf /skill-extraction-techwolf Skill Extraction with ESCO skills - TechWolf subset Dataset Summary The TECHWOLF subset, although smaller, represents a more generic distribution of job descriptions and skill spans. ESCO skills are directly annotated on the full sentence level, thus omitting the intermediate span identification step. ESCO v1.1.0 is used. This dataset is part of a three-part evaluation dataset for skill extraction: skill-extraction-tech skill-extraction-house skill-extraction-techwolf… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-techwolf.texttext-classificationn<1K3 likes218 downloads2y agoHugging Face04Cleanlab /insurance-claims-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here: https://github.com/cleanlab/structured-output-benchmark/ textn<1K1 likes166 downloads10mo agoHugging Face05Cleanlab /fire-financial-ner-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here: https://github.com/cleanlab/structured-output-benchmark/ text1K<n<10K0 likes158 downloads10mo agoHugging Face06aglazkova /keyphrase_extraction_russianA dataset for extracting or generating keywords for scientific texts in Russian. It contains annotations of scientific articles from four domains: mathematics and computer science (a fragment of the Keyphrases CS&Math Russian dataset), history, linguistics, and medicine. For more details about the dataset please refer the original paper: https://arxiv.org/abs/2409.10640 Dataset Structure abstract, an abstract in a string format; keywords, a list of keyphrases provided by… See the full description on the dataset page: https://huggingface.co/datasets/aglazkova/keyphrase_extraction_russian.texttext-generation10K<n<100K5 likes124 downloads1y agoHugging Face07Cleanlab /pii-extractionThis dataset can be used for benchmarking LLM Structured Outputs via the code here: https://github.com/cleanlab/structured-output-benchmark/ textn<1K0 likes115 downloads10mo agoHugging Face08ClarusC64 /autonomous-driving-intention-field-extraction-v0.1What this dataset tests Whether a system can infer agent intentions from context cues in complex driving scenes. This is not trajectory prediction. It is intention inference. Required outputs agent_id inferred_intention intention_confidence time_horizon_s alternative_intentions stability_score Scoring conventions confidence and stability range 0 to 1 time horizon is seconds into the near future Use case Layer one of Intention Field and Social Coherence Maps. This enables… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-intention-field-extraction-v0.1.tabulartabular-classificationn<1K0 likes105 downloads8mo agoHugging Face09Effyis /Table-Extraction Table Extract Dataset This dataset is designed to evaluate the ability of large language models (LLMs) to extract tables from text. It provides a collection of text snippets containing tables and their corresponding structured representations in JSON format. Source The dataset is based on the Table Fact Dataset, also known as TabFact, which contains 16,573 tables extracted from Wikipedia. Schema: Each data point in the dataset consists of two elements:… See the full description on the dataset page: https://huggingface.co/datasets/Effyis/Table-Extraction.textfeature-extraction10K<n<100K10 likes88 downloads2y agoHugging Face10JobOpportunitiesAPI /joa-extraction-arena JOA Job-Posting Extraction Arena Disclosure: Job Opportunities API (JOA, jobopportunitiesapi.org) is an independent data business that sells API access to job-posting data. AI helped run the experiments, check the numbers and draft this text; Loukas (Luca) Tzekos is editorially responsible. Contact: hello@jobopportunitiesapi.org. A benchmark for one narrow task: reading a real job posting and filling 11 structured fields. Job Opportunities API (JOA) built it to choose and train… See the full description on the dataset page: https://huggingface.co/datasets/JobOpportunitiesAPI/joa-extraction-arena.tabulartext-generationn<1K0 likes81 downloads2d agoHugging Face11ClarusC64 /clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1What this dataset tests Whether a model can extract the minimal cross-system coherence factor setthat explains resilience or vulnerability. It rewards minimal factor selection correct coupling recognition ranking by dominance Coherence factor labels buffering_capacity_high buffering_capacity_low variance_damping_high variance_damping_low autonomic_inflammatory_coupling sleep_metabolic_coupling stress_inflammation_coupling immune_metabolic_instability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1.texttext-classificationn<1K0 likes51 downloads8mo agoHugging Face12oberbics /Multilingual_Topic-Specific_Article-Extraction_and_Classification Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text. Cite the Dataset Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.texttext-classificationn<1K1 likes49 downloads2y agoHugging Face13ganchengguang /Text-Classification-and-Relation-Event-Extraction-Mix-datasetsThe paper of GIELLM dataset. https://arxiv.org/abs/2311.06838 Cite: @article{gan2023giellm, title={Giellm: Japanese general information extraction large language model utilizing mutual reinforcement effect}, author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori}, journal={arXiv preprint arXiv:2311.06838}, year={2023} } The dataset constructed base in livedoor news corpus 関口宏司 https://www.rondhuit.com/download.html texttext-classification1K<n<10K1 likes31 downloads2y agoHugging Face14jonathansuru /customer_service_information_extractiontextfeature-extractionn<1K4 likes24 downloads3y agoHugging Face15ClarusC64 /clinical-systemic-drift-fingerprint-extraction-v0.1What this dataset tests The whole-system drift patterninduced by a single drugover time. Required outputs drift vector affected system axes temporal profile reversibility net coherence change Use case Foundation layer for the Drug-Induced System Drift Library. texttext-classificationn<1K0 likes21 downloads8mo agoHugging Face16Balki16 /skill-extraction-tech Skill Extraction with ESCO skills - TECH subset Dataset Summary This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0). This dataset is part of a three-part evaluation dataset for skill extraction: skill-extraction-tech skill-extraction-house skill-extraction-techwolf Citation Information If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/Balki16/skill-extraction-tech.texttext-classification1K<n<10K0 likes21 downloads7mo agoHugging Face17lengocquangLAB /vie-news-tags-extractiontext10K<n<100K1 likes16 downloads2y agoHugging Face18SBB /page_extraction_dataset Title Page Extraction Dataset Description In digitised cultural heritage items such as books, newspapers and archival records, a problem that can negatively affect OCR are black margins around a page caused by document scanning. In order to enable document layout analysis (DLA), these black margins need to be cropped and the pages need to be extracted correctly. To enable the training of a machine learning model capable of extracting pages, a dataset was created. The… See the full description on the dataset page: https://huggingface.co/datasets/SBB/page_extraction_dataset.textimage-segmentation1K<n<10K0 likes16 downloads9mo agoHugging Face19KnutJaegersberg /FEVER_claim_extractionI found this dataset on my harddrive, which if I remember correctly I got from the source mentioned in the paper: "Claim extraction from text using transfer learning" - By Acharya Ashish Prabhakar, Salar Mohtaj, Sebastian Möller https://aclanthology.org/2020.icon-main.39/ The github repo with the data seems down. It extends FEVER dataset with non-claims for training claim detectors. tabular100K<n<1M0 likes14 downloads4y agoHugging Face20bala1524 /Drug_Combination_Extractiontextquestion-answering1K<n<10K1 likes14 downloads3y agoHugging Face21Tasfiya025 /Biomedical_Entity_Relation_Extractiontextn<1K0 likes13 downloads10mo agoHugging Face22ClarusC64 /clinical-crisis-invariant-signature-minimal-set-extraction-v0.1What this dataset tests Whether a model can extract the smallest cross-scale signature setthat marks the onset of a clinical crisis. It penalizes long feature lists. Outputs signature_set_top3 dominance_rank_order mechanism_hypothesis_1_sentence Signature component labels coupling_direction_flip variance_jump_no_threshold medication_response_mismatch narrative_alarm_language_shift lab_lag_reversal_pattern vitals_labs_decoupling stealth_hypoperfusion_pattern… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-crisis-invariant-signature-minimal-set-extraction-v0.1.texttext-classificationn<1K0 likes13 downloads8mo agoHugging Face23DehydratedWater42 /semantic_relations_extractiongated Dataset Card for "Semantic Relations Extraction" Dataset Description Repository The "Semantic Relations Extraction" dataset is hosted on the Hugging Face platform, and was created with code from this GitHub repository. Purpose The "Semantic Relations Extraction" dataset was created for the purpose of fine-tuning smaller LLama2 (7B) models to speed up and reduce the costs of extracting semantic relations between entities in texts. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/DehydratedWater42/semantic_relations_extraction.textsummarization10K<n<100K4 likes9 downloads3y agoHugging Face24farahbs /digit_extraction Dataset Card for Dataset Name Dataset Details Dataset Description This dataset is designed to extract and record individual digits from numbers. For each number, it contains descriptions of the digit being extracted, such as the 1st digit, 2nd digit, etc., and the corresponding extracted digit. This dataset can be used for tasks such as training digit extraction models, digit recognition, and sequence-based learning tasks. Curated by: BEN SLAMA Farah… See the full description on the dataset page: https://huggingface.co/datasets/farahbs/digit_extraction.text10K<n<100K0 likes9 downloads2y agoHugging Face25JesseZz /event_extractiontext1K<n<10K0 likes6 downloads2y agoHugging Face26Savindu9x /clinical-entity-extraction-transformedtext10K<n<100K0 likes4 downloads1y agoHugging Face27Orib16 /main_product_extractiontextn<1K0 likes3 downloads3y agoHugging Face28Appz7 /t5-metadata-extraction-dataset1textn<1K0 likes3 downloads2y agoHugging Face29Appz7 /t5-metadata-extraction-dataset2textn<1K0 likes3 downloads2y agoHugging Face30ejung /Triplets-Extraction2textn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.