datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
events_classification_biotech
Key aspects
Event extraction;
Multi-label classification;
Biotech news domain;
31 classes;
3140 total number of examples;
Motivation
Text classification is a widespread task and a foundational step in numerous information extraction pipelines. However, a notable challenge in current NLP research lies in the oversimplification of benchmarking datasets, which predominantly focus on rudimentary tasks such as topic classification or sentiment analysis.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/events_classification_biotech.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.cycic_classificationhttps://storage.googleapis.com/ai2-mosaic/public/cycic/CycIC-train-dev.zip
https://colab.research.google.com/drive/16nyxZPS7-ZDFwp7tn_q72Jxyv0dzK1MP?usp=sharing
@article{Kejriwal2020DoFC,
title={Do Fine-tuned Commonsense Language Models Really Generalize?},
author={Mayank Kejriwal and Ke Shen},
journal={ArXiv},
year={2020},
volume={abs/2011.09159}
}
added for
@article{sileo2023tasksource,
title={tasksource: Structured Dataset Preprocessing Annotations for Frictionless Extreme… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/cycic_classification.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
it-support-l1-ticket-classification
IT Support L1 Multilingual Dataset
Dataset Summary
IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping.
This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1_complete
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
Patent_classification_QAVietnamese-toxic-classificationQuestion-Classification-With-AnswersVietnamese-toxic-classificationText_classification_by_subject_area
🇰🇿 Kazakh Topic and Domain Identification Dataset
Dataset Summary
Kazakh Topic and Domain Identification Dataset is a Kazakh-language instruction-following dataset designed for topic recognition, domain classification, and text understanding tasks.
Each sample contains a short Kazakh prompt, a long Kazakh text passage, a target response, a domain label, and a unique sample identifier. The dataset is intended to help Large Language Models (LLMs) and NLP systems… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Text_classification_by_subject_area.prompts-classification-pfg
