Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simana /textclassificationMNLItext100K<n<1M0 likes351 downloads4y agoHugging Face02CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes349 downloads5mo agoHugging Face03owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K31 likes275 downloads4y agoHugging Face04knowledgator /Scientific-text-classificationtext10K<n<100K21 likes222 downloads3y agoHugging Face05israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes76 downloads5y agoHugging Face06ml4pubmed /pubmed-text-classification-cased ml4pubmed/pubmed-text-classification-cased A parsed/cleaned version of the source data retaining case. texttext-classification1M<n<10M0 likes51 downloads4y agoHugging Face07ganchengguang /Text-Classification-and-Relation-Event-Extraction-Mix-datasetsThe paper of GIELLM dataset. https://arxiv.org/abs/2311.06838 Cite: @article{gan2023giellm, title={Giellm: Japanese general information extraction large language model utilizing mutual reinforcement effect}, author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori}, journal={arXiv preprint arXiv:2311.06838}, year={2023} } The dataset constructed base in livedoor news corpus 関口宏司 https://www.rondhuit.com/download.html texttext-classification1K<n<10K1 likes40 downloads2y agoHugging Face08Veer15 /cancer-text-classificationtext1K<n<10K1 likes39 downloads3y agoHugging Face09jakeazcona /short-text-multi-labeled-emotion-classificationtabular10K<n<100K2 likes26 downloads5y agoHugging Face10sobamchan /ja-toxic-text-classification-open2ch Open 2ch-based toxic classification dataset Based on p1atdev/open2ch We apply keyword-based filtering to collect toxic texts We use Perspective API to filter non-toxic texts from the original corpus 3k texts for each class, toxic (label=1) and non-toxic (label=0) texts perspective_api_score is a prediction of toxicity score by the Perspective API tabular1K<n<10K1 likes26 downloads2y agoHugging Face11GuillermoTBB /charles-dickens-text-classification Dataset Description This Dataset is designed to classify paragraphs of text as either written by Charles Dickens or generated to imitate various distinct writing styles. The primary use case is in the domain of literary analysis and text generation, where it can help distinguish between authentic Dickensian text and stylistic imitations. The dataset was created from paragraphs extracted from “Great Expectations” by Charles Dickens. The original text was taken from the Gutenberg… See the full description on the dataset page: https://huggingface.co/datasets/GuillermoTBB/charles-dickens-text-classification.texttext-classification1K<n<10K1 likes23 downloads2y agoHugging Face12AkashPrasadMishra /Hierarchical_Text_Classification_Intent_Classificationtexttext-classification1K<n<10K0 likes19 downloads2y agoHugging Face13sanjeettoosi /multi-label-text-classificationtext1K<n<10K0 likes18 downloads2y agoHugging Face14Nerdy37 /ai-human-text-classification AI vs Human Sentence Classification Dataset Dataset Summary sentence_dataset is a sentence-level binary classification dataset containing approximately 9.84 million sentences labelled as either AI-generated (1) or human-written (0). It was constructed by extracting individual sentences from two source datasets and merging them: Dataset 1 — ai_vs_human_content_v2_20000.csv: 20,000 rows of short text and code snippets with rich metadata (prompt, topic, source… See the full description on the dataset page: https://huggingface.co/datasets/Nerdy37/ai-human-text-classification.tabulartext-classification1M<n<10M0 likes17 downloads4mo agoHugging Face15omkar56 /text_category_classificationtextn<1K2 likes16 downloads3y agoHugging Face16violetakastreva /line-level-code-vs-text-classificationThis dataset was created for SemEval-2026 Task 13, which focuses on distinguishing machine-generated code from human-written code across multiple programming languages and domains. While the original SemEval task operates at the code snippet level, this dataset provides line-level annotations that enable finer-grained analysis of how code-like and text-like content is distributed within mixed inputs. The dataset is intended to support research in machine-generated code detection, robust… See the full description on the dataset page: https://huggingface.co/datasets/violetakastreva/line-level-code-vs-text-classification.texttext-classification10K<n<100K0 likes14 downloads8mo agoHugging Face17haoxianc /kalashnikov1405_facebook-text-classification Facebook Text classification Mirror of the Kaggle dataset kalashnikov1405/facebook-text-classification by kalashnikov1405, released under CC0: Public Domain. All credit goes to the original author; please cite and link the Kaggle page when using this data. Facebook text classification of the dataset of 5000 row Original description (from Kaggle) The Facebook Text Classification Dataset consists of 5,000 social media posts designed for text analytics and machine… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/kalashnikov1405_facebook-text-classification.tabular1K<n<10K0 likes13 downloads6h agoHugging Face18Mahadih534 /arabic-text-classificationtext100K<n<1M1 likes11 downloads2y agoHugging Face19McOwska /action-required-text-classificationThis dataset was created to support the task of classifying short text messages based on whether they require user action. The goal is to distinguish between different types of communicative intent, specifically whether a message requires immediate action, suggests optional action, or is purely informational. The dataset consists of short sentences resembling real-world messages, such as system notifications, emails, or app prompts. Each sample is labeled with one of three classes:… See the full description on the dataset page: https://huggingface.co/datasets/McOwska/action-required-text-classification.textn<1K1 likes10 downloads5mo agoHugging Face20tor24 /short-text-classification-v1 Short Text Classification v1 A small human-written dataset of short texts labeled by sentiment and risk perception. Intended Use Text classification, sentiment analysis, and educational research. License CC-BY-4.0 textn<1K0 likes7 downloads9mo agoHugging Face21windcrossroad /text-classification-chatGPT-100textn<1K1 likes6 downloads3y agoHugging Face22krushilpatel /covid-tweet-text-classificationtabular1K<n<10K0 likes6 downloads3y agoHugging Face23amruta333 /text_classificationtexttext-classification1K<n<10K0 likes5 downloads3y agoHugging Face24Rickie111 /text_category_classificationtextn<1K0 likes5 downloads5mo agoHugging Face25AmruthaS /TextClassificationtextn<1K1 likes4 downloads3y agoHugging Face26mishrasaurabh847 /covid-tweet-text-classificationtabular10K<n<100K0 likes4 downloads3y agoHugging Face27comitium /writing-models-text-classification-6006738d_300_3000text1K<n<10K0 likes4 downloads1y agoHugging Face28tommybrenson /text-classificationtabular1K<n<10K0 likes3 downloads2y agoHugging Face29MMEX /text-classification-datasettextn<1K0 likes3 downloads2y agoHugging Face30comitium /writing-models-text-classification-a5c5b635text1K<n<10K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.