Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Karavet /ILUR-news-text-classification-corpus News Texts Dataset We release a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society. The articles are split into train (2242k tokens) and test sets (425k tokens). For more details, refer to the paper. texttext-classification100K<n<1M7 likes1.3k downloads4y agoHugging Face02ourafla /Mental-Health_Text-Classification_Dataset Mental Health Text Classification Dataset (4-Class) Dataset Description This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models. The repository includes: An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.texttext-classification10K<n<100K13 likes763 downloads10mo agoHugging Face03simana /textclassificationMNLItext100K<n<1M0 likes351 downloads4y agoHugging Face04CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes349 downloads5mo agoHugging Face05Lots-of-LoRAs /task679_hope_edi_english_text_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task679_hope_edi_english_text_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task679_hope_edi_english_text_classification.texttext-generation1K<n<10K0 likes324 downloads2y agoHugging Face06Rami /multi-label-class-github-issues-text-classification Dataset Card for "multi-label-class-github-issues-text-classification" More Information needed text1K<n<10K2 likes315 downloads4y agoHugging Face07owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K31 likes275 downloads4y agoHugging Face08argilla /end2end_textclassification Dataset Card for end2end_textclassification This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/argilla/end2end_textclassification.text1K<n<10K2 likes270 downloads2y agoHugging Face09jakeazcona /short-text-labeled-emotion-classificationtext10K<n<100K5 likes262 downloads5y agoHugging Face10knowledgator /Scientific-text-classificationtext10K<n<100K21 likes222 downloads3y agoHugging Face11murodbek /uz-text-classification Dataset Card for "uzbek_news" Dataset Summary Multi-label text classification dataset for Uzbek language and some sourcode for analysis. This repository contains the code and dataset used for text classification analysis for the Uzbek language. The dataset consists text data from 9 Uzbek news websites and press portals that included news articles and press releases. These websites were selected to cover various categories such as politics, sports, entertainment… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uz-text-classification.texttext-classification100K<n<1M7 likes206 downloads3y agoHugging Face12ViravirastSHZ /Hafez-text-classification-datasettext10K<n<100K0 likes204 downloads2y agoHugging Face13reubenjohn /stackoverflow-unified-text-open-status-classification Dataset Card for "stackoverflow-unified-text-open-status-classification" More Information needed tabular1M<n<10M0 likes175 downloads4y agoHugging Face14Lots-of-LoRAs /task1645_medical_question_pair_dataset_text_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1645_medical_question_pair_dataset_text_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1645_medical_question_pair_dataset_text_classification.texttext-generation1K<n<10K2 likes152 downloads2y agoHugging Face15hossein20s /enrun-emails-text-classificationtext10K<n<100K5 likes131 downloads4y agoHugging Face16nickmuchi /financial-text-combo-classification Dataset Card for "financial-text-combo-classification" More Information needed texttext-classification10K<n<100K15 likes131 downloads4y agoHugging Face17cestwc /text_classificationtext1M<n<10M0 likes130 downloads3y agoHugging Face18marcelsun /wos_hierarchical_multi_label_text_classificationIntroduced by du Toit and Dunaiski (2024) Introducing Three New Benchmark Datasets for Hierarchical Text Classification. The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created… See the full description on the dataset page: https://huggingface.co/datasets/marcelsun/wos_hierarchical_multi_label_text_classification.texttext-classification100K<n<1M0 likes122 downloads2y agoHugging Face19argilla /synthetic-text-classification-news Dataset Card for synthetic-text-classification-news This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/argilla/synthetic-text-classification-news/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news.textn<1K17 likes115 downloads2y agoHugging Face20argilla /synthetic-text-classification-news-multi-label Dataset Card for synthetic-text-classification-news-multi-label This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/synthetic-text-classification-news-multi-label/raw/main/pipeline.yaml" or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news-multi-label.textn<1K8 likes111 downloads2y agoHugging Face21argilla /synthetic-domain-text-classification Dataset Card for my-distiset-b845cf19 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/my-distiset-b845cf19/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-domain-text-classification.texttext-classification1K<n<10K10 likes110 downloads2y agoHugging Face22MCINext /synthetic-persian-text-keyword-pair-classification Dataset Summary Synthetic Persian Text-Keywords Pair Classification (SynPerTextKeywordsPC) is a Persian (Farsi) dataset developed for the Pair Classification task. The dataset focuses on determining whether a keyword or short phrase is relevant to a longer Persian text passage. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically created using GPT-4o-mini. Language(s): Persian (Farsi) Task(s): Pair Classification (Text–Keyword Relevance)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-text-keyword-pair-classification.text10K<n<100K0 likes101 downloads1y agoHugging Face23Lots-of-LoRAs /task1605_ethos_text_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1605_ethos_text_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1605_ethos_text_classification.texttext-generationn<1K0 likes86 downloads2y agoHugging Face24sdiazlor /text-classification-news-topics Dataset Card for test This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/test/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/text-classification-news-topics.text1K<n<10K1 likes86 downloads2y agoHugging Face25Lots-of-LoRAs /task1607_ethos_text_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1607_ethos_text_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1607_ethos_text_classification.texttext-generationn<1K0 likes80 downloads2y agoHugging Face26israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes76 downloads5y agoHugging Face27MCINext /synthetic-persian-text-tone-classification-v3 Dataset Summary Synthetic Persian Text Tone Classification (SynPerTextToneClassification) - Version 2 is a Persian (Farsi) dataset made for the Classification task. It focuses on figuring out the tonal content of text and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was created synthetically using the GPT-4o-mini model, offering examples across two main tones: عامیانه (informal/colloquial) and رسمی (formal). Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-text-tone-classification-v3.text10K<n<100K0 likes75 downloads1y agoHugging Face28argilla /end2end_textclassification_with_suggestions_and_responses Dataset Card for end2end_textclassification_with_suggestions_and_responses This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure… See the full description on the dataset page: https://huggingface.co/datasets/argilla/end2end_textclassification_with_suggestions_and_responses.text1K<n<10K5 likes74 downloads2y agoHugging Face29Lots-of-LoRAs /task1606_ethos_text_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1606_ethos_text_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1606_ethos_text_classification.texttext-generationn<1K0 likes73 downloads2y agoHugging Face30Lots-of-LoRAs /task1341_msr_text_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1341_msr_text_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1341_msr_text_classification.texttext-generationn<1K0 likes72 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.