datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jigsaw-toxic-comment-classification-challenge
Dataset Description
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
You must create a model which predicts a probability of each type of toxicity for each comment.
File descriptions
train.csv - the training set, contains comments with their binary labels
test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.jailbreak-classification
Jailbreak Classification
Dataset Summary
Dataset used to classify prompts as jailbreak vs. benign.
Dataset Structure
Data Fields
prompt: an LLM prompt
type: classification label, either jailbreak or benign
Dataset Creation
Curation Rationale
Created to help detect & prevent harmful jailbreak prompts when users interact with LLMs.
Source Data
Jailbreak prompts sourced from: https://github.com/verazuo/jailbreak_llms… See the full description on the dataset page: https://huggingface.co/datasets/jackhhao/jailbreak-classification.commit-classification-dataset
Commit Classification Dataset
This dataset is designed for multi-label classification of Git commit messages into predefined categories.
Dataset Summary
This dataset contains:
Training data: Commit messages and their corresponding labels for training the model.
Validation data: A separate set of messages for tuning and evaluation.
Testing data: Unlabeled commit messages for testing the model’s performance.
The goal of the dataset is to classify each commit message into… See the full description on the dataset page: https://huggingface.co/datasets/meriemm6/commit-classification-dataset.news-political-bias-classification-datasetDataset actually from kaggle.
Couldn't find it here so I uploaded it.
commit-classification-17ktoxicity_classification_jigsaw
Dataset info
Training Dataset:
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
The original dataset can be found here: jigsaw_toxic_classification
Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes.
Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.NyayaAnumana-Classification-Datasms-spam-classificationtabular-benchmark-797-classificationPubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.adaption-defi-wallet-risk-classification-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-defi_wallet_risk_classification
This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks based on behavioral features. Each sample provides wallet metrics such as transaction counts, action ratios, and concentration levels, followed by a binary risk label and a concise reasoning statement. The data is designed for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification-v1.Scientific-text-classificationcancer_data_classificationphishing_url_classification
Phishing URL Classification Dataset
This dataset contains URLs labeled as 'Safe' (0) or 'Not Safe' (1) for phishing detection tasks.
Dataset Summary
This dataset contains URLs labeled for phishing detection tasks. It's designed to help train and evaluate models that can identify potentially malicious URLs.
Dataset Creation
The dataset was synthetically generated using a custom script that creates both legitimate and potentially phishing URLs. This approach… See the full description on the dataset page: https://huggingface.co/datasets/imanoop7/phishing_url_classification.cis5300-text-classification
Complex Word Identification (CIS 5300)
Dataset Description
This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple.
CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.International_Classification_Diseases_Clinical_Modification_icd10cm_order_April_2024ahsan81_hotel-reservations-classification-dataset
Hotel Reservations Dataset
Can you predict if customer is going to cancel the reservation ?
Dataset Info
Source: Kaggle
Original Size: 0.47 MB
Kaggle Downloads: 57,080
Files: 1
Files
Hotel Reservations.csv
Mirrored from Kaggle
typhoon-intensity-classification
Typhoon - Image Classification Dataset
This dataset comes from PTIT AI Challenge and is organized for a multi-class image classification task focusing on tropical cyclone (typhoon) intensity estimation.
Dataset Structure
The directory structure is organized as follows:
train/
├── images/
│ ├── image1.jpg
│ └── ...
└── annotations.csv (only present in the train folder)
The public_test and private_test sets are used to evaluate and score the… See the full description on the dataset page: https://huggingface.co/datasets/star092304/typhoon-intensity-classification.clickbait_title_classificationDataset introduced in Stop Clickbait: Detecting and Preventing Clickbaits in Online News Mediaby Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, Niloy Ganguly
Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, and Niloy Ganguly. "Stop Clickbait: Detecting and Preventing Clickbaits in Online News Media”. In Proceedings of the 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), San Fransisco, US, August 2016.
Cite:… See the full description on the dataset page: https://huggingface.co/datasets/marksverdhei/clickbait_title_classification.japanese-image-classification-evaluation-dataset
recruit-jp/japanese-image-classification-evaluation-dataset
Overview
Developed by: Recruit Co., Ltd.
Dataset type: Image Classification
Language(s): Japanese
LICENSE: CC-BY-4.0
More details are described in our tech blog post.
日本語CLIP学習済みモデルとその評価用データセットの公開
Dataset Details
This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks.
jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.Near-Earth-Comets-Classification-dataset
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me
Near-Earth Comets (NECs) Classification Dataset
This repository provides an open-source datasetof Near‑Earth Comets
(NECs) and their classification as Potentially Hazardous… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/Near-Earth-Comets-Classification-dataset.patent_classificationurl-classifications
Model Card: URL Classifications Dataset
Dataset Summary
The URL Classifications Dataset is a collection of URL classifications for PDF documents, primarily derived from the SafeDocs corpus. It contains multiple CSV files with different subsets of classifications, including both raw and processed data.
Supported Tasks
This dataset supports the following tasks:
Text Classification
URL-based Document Classification
PDF Content Inference
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/snats/url-classifications.cancer_classificationemail-spam-classification
Email Spam Classification
The dataset consists of a collection of emails categorized into two major classes: spam and not spam. It is designed to facilitate the development and evaluation of spam detection or email filtering systems.
The spam emails in the dataset are typically unsolicited and unwanted messages that aim to promote products or services, spread malware, or deceive recipients for various malicious purposes. These emails often contain misleading subject lines… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/email-spam-classification.Amharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.malayalam_news_classificationpubmed-classification-20k
ml4pubmed/pubmed-classification-20k
20k subset of pubmed text classification from course
gandpablo_news-articles-for-political-bias-classification
News articles for political bias classification
Mirror of the Kaggle dataset gandpablo/news-articles-for-political-bias-classification by Pablo Gandia, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data.
Text from 10k+ English news articles classified by political bias
License
MIT License, Copyright (c) Pablo Gandia. The full license text is in LICENSE and applies to all files in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/gandpablo_news-articles-for-political-bias-classification.KOTOX-classification
KOTOX
: A Korean Toxic Dataset for Deobfuscation and Detoxification
Hate Speech Detection dataset 👉 Here!Detoxification or Sanitization dataset 👉 KOTOX
📚 paper |
🐈⬛ git
📝 Dataset Summary
KOTOX is the first Korean dataset designed for both toxic text detoxification and obfuscation robustness.
It provides paired neutral-toxic sentences and their obfuscated counterparts, constructed with 17 linguistically grounded transformation rules reflecting the… See the full description on the dataset page: https://huggingface.co/datasets/ssgyejin/KOTOX-classification.
