Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thesofakillers /jigsaw-toxic-comment-classification-challenge Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate You must create a model which predicts a probability of each type of toxicity for each comment. File descriptions train.csv - the training set, contains comments with their binary labels test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular100K<n<1M13 likes9.7k downloads2y agoHugging Face02Arsive /toxicity_classification_jigsaw Dataset info Training Dataset: You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate The original dataset can be found here: jigsaw_toxic_classification Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes. Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.tabulartext-classification100K<n<1M5 likes560 downloads3y agoHugging Face03owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K33 likes272 downloads4y agoHugging Face04imodels /tabular-benchmark-797-classificationtabular1K<n<10K0 likes271 downloads3y agoHugging Face05TachyHealth /International_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024tabular10K<n<100K6 likes163 downloads3y agoHugging Face06jason1966 /ahsan81_hotel-reservations-classification-dataset Hotel Reservations Dataset Can you predict if customer is going to cancel the reservation ? Dataset Info Source: Kaggle Original Size: 0.47 MB Kaggle Downloads: 57,080 Files: 1 Files Hotel Reservations.csv Mirrored from Kaggle tabular10K<n<100K0 likes155 downloads6mo agoHugging Face07seniruk /Near-Earth-Comets-Classification-dataset Hi, I’m Seniru Epasinghe 👋 I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier. 🌐 Connect with me          Near-Earth Comets (NECs) Classification Dataset This repository provides an open-source datasetof Near‑Earth Comets (NECs) and their classification as Potentially Hazardous… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/Near-Earth-Comets-Classification-dataset.tabular100K<n<1M1 likes149 downloads11mo agoHugging Face08CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes130 downloads5mo agoHugging Face09israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes118 downloads5y agoHugging Face10snats /url-classifications Model Card: URL Classifications Dataset Dataset Summary The URL Classifications Dataset is a collection of URL classifications for PDF documents, primarily derived from the SafeDocs corpus. It contains multiple CSV files with different subsets of classifications, including both raw and processed data. Supported Tasks This dataset supports the following tasks: Text Classification URL-based Document Classification PDF Content Inference Languages The… See the full description on the dataset page: https://huggingface.co/datasets/snats/url-classifications.tabular1M<n<10M12 likes118 downloads2y agoHugging Face11tussiiiii /llm-classification-distilled-v2-sharded LLM Classification Distilled v2 Sharded Overview This repository stores shard CSV files produced by the teacher-judge distillation pipeline. How to Use Run the distillation notebook once per shard: NUM_SHARDS = 4 SHARD_INDEX = 0 .. 3 After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos. Final Repositories Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.tabulartext-classification100K<n<1M0 likes75 downloads5mo agoHugging Face12Puidii /trustpilot_review_classification Github Repo: https://github.com/pattplatt/trustpilot-review-classification Project report: https://homepageblob.blob.core.windows.net/content/Using NLP Techniques to Infer Trustpilot Ratings from User Reviews.pdf tabulartext-classification100K<n<1M1 likes60 downloads2y agoHugging Face13noanabeshima /forecastability_classificationThis dataset is composed of Claude-labelled fineweb documents. For each document, Claude is asked if it is 'forecastable' (i.e. would be a reasonable seed for a pastcasting question) and to estimate the date the document was published. V1 splits were generated by having Claude label ~50K random fineweb documents and v2 splits were augmented with labels on ~30K additional documents that a DebertaV3 classifier finetuned on ratio10_v1 thought were forecastable (Claude thought ~1/3 of these… See the full description on the dataset page: https://huggingface.co/datasets/noanabeshima/forecastability_classification.tabular100K<n<1M0 likes47 downloads1y agoHugging Face14hajili /azerbaijani_review_sentiment_classificationAzerbaijani Sentiment Classification Dataset with ~160K reviews. Dataset contains 3 columns: Content, Score, Upvotes tabulartext-classification100K<n<1M6 likes44 downloads3y agoHugging Face15hotal /emergency_classification Emergency Messages Classification Dataset tabulartext-classification10K<n<100K1 likes41 downloads3y agoHugging Face16facells /chronos-historical-dataset-sdt-phase-classificationThe schema o fthe dataset is the following: polityid: a Polity ID formatted with a standard method: 2 letters to indicate the area of origin of the culture, 3 letters to indicate the name of the polity, 1 letter to indicate the type of society (c=culture/community; n=nomads; e=empire; k=kingdom; r=republic) and 1 letter to indicate the periodization (t=terminal; l=late; m=middle; e=early; f=formative; i=initial; =any). For example “EsSpael” is the late Spanish Empire, “ItRomre” is the early… See the full description on the dataset page: https://huggingface.co/datasets/facells/chronos-historical-dataset-sdt-phase-classification.tabulartext-classification1K<n<10K0 likes41 downloads2y agoHugging Face17CryptoTaxEdge /crypto-tax-classification-2026h1 Crypto tax transaction classification, 2026 H1 labelled benchmark 388 public blockchain transactions on Ethereum, Polygon, and Arbitrum, each labelled with a transaction category and the default US tax treatment that category maps to under an open, versioned classification schema. Published under CC BY 4.0 together with the schema itself. Canonical landing page: https://cryptotaxedge.com/research/benchmarks/2026-h1/ Schema: https://cryptotaxedge.com/standard/ DOI:… See the full description on the dataset page: https://huggingface.co/datasets/CryptoTaxEdge/crypto-tax-classification-2026h1.texttext-classificationn<1K0 likes41 downloads2mo agoHugging Face18Farmaanaa /iran_hs_customs_classification طبقه‌بندیِ کالا (HS) — تعرفهٔ گمرک ایران جدولِ طبقه‌بندیِ نظامِ هماهنگ‌شدهٔ توصیف و کدگذاری کالا (HS) برای تعرفهٔ گمرک ایران، با سلسله‌مراتبِ کامل: فصل (۲ رقم) ← عنوان (۴) ← زیرعنوان (۶، بین‌المللی) ← تعرفه (۸ رقم، ملی). ستون توضیح code کدِ HS level فصل / عنوان / زیرعنوان / تعرفه digits تعداد رقم (۲/۴/۶/۸) name_fa شرحِ فارسیِ کالا parent_code کدِ والد (برای تجمیع) ۱۴٬۷۳۷ کد شامل ۸٬۴۰۰ ردیفِ تعرفهٔ هشت‌رقمی. مرجعی برای پیوستن به داده‌های تجارت و تحلیلِ… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/iran_hs_customs_classification.tabular10K<n<100K0 likes39 downloads17d agoHugging Face19NLPC-UOM /Sinhala-News-Source-classificationThis dataset contains Sinhala news headlines extracted from 9 news sources (websites) (Sri Lanka Army, Dinamina, GossipLanka, Hiru, ITN, Lankapuwath, NewsLK, Newsfirst, World Socialist Web Site-Sinhala). This is a processed version of the corpus created by Sachintha, D., Piyarathna, L., Rajitha, C., and Ranathunga, S. (2021). Exploiting parallel corpora to improve multilingual embedding based document and sentence alignment. Single word sentences, invalid characters have been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Source-classification.tabulartext-classification10K<n<100K0 likes38 downloads4y agoHugging Face20noanabeshima /forecastability_classification_oldtabular100K<n<1M0 likes37 downloads1y agoHugging Face21jakeazcona /short-text-multi-labeled-emotion-classificationtabular10K<n<100K2 likes34 downloads5y agoHugging Face22haoxianc /kalashnikov1405_facebook-text-classification Facebook Text classification Mirror of the Kaggle dataset kalashnikov1405/facebook-text-classification by kalashnikov1405, released under CC0: Public Domain. All credit goes to the original author; please cite and link the Kaggle page when using this data. Facebook text classification of the dataset of 5000 row Original description (from Kaggle) The Facebook Text Classification Dataset consists of 5,000 social media posts designed for text analytics and machine… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/kalashnikov1405_facebook-text-classification.tabular1K<n<10K0 likes34 downloads4d agoHugging Face23llangnickel /long-covid-classification-data Data Description Long-COVID related articles have been manually collected by information specialists.Please find further information here. Size Training Development Test Total Positive Examples 215 76 70 345 Negative Examples 199 62 68 345 Total 414 238 138 690 Citation @article{10.1093/database/baac048,author = {Langnickel, Lisa and Darms, Johannes and Heldt, Katharina and Ducks, Denise and Fluck, Juliane},title = "{Continuous development… See the full description on the dataset page: https://huggingface.co/datasets/llangnickel/long-covid-classification-data.tabulartext-classificationn<1K2 likes31 downloads4y agoHugging Face24phucthaiv02 /Jigsaw-Agile-Community-Rules-Classificationtabular1K<n<10K0 likes31 downloads1y agoHugging Face25sobamchan /ja-toxic-text-classification-open2ch Open 2ch-based toxic classification dataset Based on p1atdev/open2ch We apply keyword-based filtering to collect toxic texts We use Perspective API to filter non-toxic texts from the original corpus 3k texts for each class, toxic (label=1) and non-toxic (label=0) texts perspective_api_score is a prediction of toxicity score by the Perspective API tabular1K<n<10K1 likes30 downloads2y agoHugging Face26haoxianc /lethaldiran_fakenofake-classification-using-stopfake-source Fake/No-Fake classification using Stopfake source Mirror of the Kaggle dataset lethaldiran/fakenofake-classification-using-stopfake-source by lethaldiran, released under CC0: Public Domain. All credit goes to the original author; please cite and link the Kaggle page when using this data. stopfake + western news Original description (from Kaggle) No description provided. tabularn<1K0 likes30 downloads4d agoHugging Face27jonaskoenig /topic_classificationtabular10M<n<100M1 likes27 downloads4y agoHugging Face28thomasavare /waste-classificationtabular10K<n<100K1 likes25 downloads3y agoHugging Face29rachana-rj-news-classification /headline_datatabular1K<n<10K0 likes25 downloads2y agoHugging Face30thomasavare /waste-classification-v2 Dataset Card for Dataset Name Dataset Summary Dataset used to train a language model to do classification on 50 different waste classes. Languages English Dataset Structure Data Instances Phrase Class Index "I have this apple phone charger to throw, where should I put it ?" PHONE CHARGER 26 "Should I recycle a disposable cup ?" Plastic Cup 32 "I have a milk brick" Tetrapack 45 Data Fields Phrase Class… See the full description on the dataset page: https://huggingface.co/datasets/thomasavare/waste-classification-v2.tabular10K<n<100K1 likes24 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.