Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thesofakillers /jigsaw-toxic-comment-classification-challenge Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate You must create a model which predicts a probability of each type of toxicity for each comment. File descriptions train.csv - the training set, contains comments with their binary labels test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular100K<n<1M13 likes9.8k downloads2y agoHugging Face02jackhhao /jailbreak-classification Jailbreak Classification Dataset Summary Dataset used to classify prompts as jailbreak vs. benign. Dataset Structure Data Fields prompt: an LLM prompt type: classification label, either jailbreak or benign Dataset Creation Curation Rationale Created to help detect & prevent harmful jailbreak prompts when users interact with LLMs. Source Data Jailbreak prompts sourced from: https://github.com/verazuo/jailbreak_llms… See the full description on the dataset page: https://huggingface.co/datasets/jackhhao/jailbreak-classification.texttext-classification1K<n<10K87 likes5.8k downloads3y agoHugging Face03SetFit /ade_corpus_v2_classification ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data. This is a dataset for classification if a sentence is ADE-related (True) or not (False). Train size: 17,637 Test size: 5,879 Source dataset Paper text10K<n<100K6 likes2.4k downloads4y agoHugging Face04ccdv /arxiv-classificationArxiv Classification: a classification of Arxiv Papers (11 classes). This dataset is intended for long context classification (documents have all > 4k tokens). Copied from "Long Document Classification From Local Word Glimpses via Recurrent Attention Learning" @ARTICLE{8675939, author={He, Jun and Wang, Liqun and Liu, Liu and Feng, Jiao and Wu, Hao}, journal={IEEE Access}, title={Long Document Classification From Local Word Glimpses via Recurrent Attention Learning}, year={2019}… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/arxiv-classification.texttext-classification10K<n<100K27 likes2.3k downloads2y agoHugging Face05torchsight /cybersecurity-classification-benchmark TorchSight Cybersecurity Classification Benchmark A two-tier benchmark dataset for evaluating cybersecurity document classifiers, released with the TorchSight system. Used in: Dobrovolskyi, I. Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System. Journal of Information Security and Applications, 2026. Canonical per-model numbers live in BENCHMARK_NUMBERS.md, auto-generated from the per-prediction result JSONs… See the full description on the dataset page: https://huggingface.co/datasets/torchsight/cybersecurity-classification-benchmark.texttext-classification1K<n<10K1 likes2.1k downloads5mo agoHugging Face06tattabio /ec_classificationtextn<1K0 likes1.7k downloads2y agoHugging Face07meriemm6 /commit-classification-dataset Commit Classification Dataset This dataset is designed for multi-label classification of Git commit messages into predefined categories. Dataset Summary This dataset contains: Training data: Commit messages and their corresponding labels for training the model. Validation data: A separate set of messages for tuning and evaluation. Testing data: Unlabeled commit messages for testing the model’s performance. The goal of the dataset is to classify each commit message into… See the full description on the dataset page: https://huggingface.co/datasets/meriemm6/commit-classification-dataset.texttext-classification1K<n<10K0 likes1.5k downloads2y agoHugging Face08mteb /multilingual-scala-classification ScalaClassification An MTEB dataset Massive Text Embedding Benchmark ScaLa a linguistic acceptability dataset for the mainland Scandinavian languages automatically constructed from dependency annotations in Universal Dependencies Treebanks. Published as part of 'ScandEval: A Benchmark for Scandinavian Natural Language Processing' Task category t2c Domains Fiction, News, Non-fiction, Blog, Spoken, Web, Written Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-scala-classification.texttext-classification10K<n<100K1 likes1.5k downloads8mo agoHugging Face09Kushal0532 /news-political-bias-classification-datasetDataset actually from kaggle. Couldn't find it here so I uploaded it. text10K<n<100K0 likes1.5k downloads1y agoHugging Face10amir-kazemi /aidovecl-vehicle-detection-classification-localization AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models. Citation Notice Please ensure that all publications and presentations using this data reference the following paper: Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.imageobject-detection1K<n<10K0 likes1.4k downloads6mo agoHugging Face11ccdv /patent-classificationPatent Classification: a classification of Patents and abstracts (9 classes). This dataset is intended for long context classification (non abstract documents are longer that 512 tokens). Data are sampled from "BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization." by Eva Sharma, Chen Li and Lu Wang See: https://aclanthology.org/P19-1212.pdf See: https://evasharma.github.io/bigpatent/ It contains 9 unbalanced classes, 35k Patents and abstracts divided into 3 splits:… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/patent-classification.texttext-classification10K<n<100K30 likes1.3k downloads2y agoHugging Face12Karavet /ILUR-news-text-classification-corpus News Texts Dataset We release a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society. The articles are split into train (2242k tokens) and test sets (425k tokens). For more details, refer to the paper. texttext-classification100K<n<1M7 likes1.3k downloads4y agoHugging Face13mteb /multilingual-sentiment-classification MultilingualSentimentClassification An MTEB dataset Massive Text Embedding Benchmark Sentiment classification dataset with binary (positive vs negative sentiment) labels. Includes 30 languages and dialects. Task category t2c DomainsReviews, Written Reference https://huggingface.co/datasets/mteb/multilingual-sentiment-classification How to evaluate on this task You can evaluate an embedding model on this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-sentiment-classification.texttext-classification100K<n<1M1 likes1.2k downloads1y agoHugging Face14rsh-raj /commit-classification-17ktext10K<n<100K0 likes1.2k downloads2y agoHugging Face15Lots-of-LoRAs /task903_deceptive_opinion_spam_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task903_deceptive_opinion_spam_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task903_deceptive_opinion_spam_classification.texttext-generation1K<n<10K1 likes1.1k downloads2y agoHugging Face16Lots-of-LoRAs /task902_deceptive_opinion_spam_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task902_deceptive_opinion_spam_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task902_deceptive_opinion_spam_classification.texttext-generation1K<n<10K0 likes1k downloads2y agoHugging Face17vnahata /AfriMCQA-category-classification Afri-MCQA cross-modal cultural category classification (MTEB) Classify the cultural category of an entry from its photograph and the question about it spoken by a native speaker, across 16 African languages. Labels index this list: geography, building, and landmarks public figure and pop culture cooking and food objects, materials, clothing tranditions, art, and history brands, products, and companies plants and animals people, and everyday life vehicles and transportation… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-category-classification.audioaudio-classification1K<n<10K0 likes961 downloads1mo agoHugging Face18ManuD /dfl_classification_512text100K<n<1M1 likes955 downloads4y agoHugging Face19aliencaocao /multimodal_meme_classification_singapore Dataset Card for Offensive Memes in Singapore Context Dataset Details Dataset Description This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards. Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.imagetext-generation100K<n<1M1 likes863 downloads2y agoHugging Face20dlab-spp /safety-classifications Safety Annotations for dolma3_mix Safety score annotations for a 20K-shard subset of allenai/dolma3_mix-6T using locuslab/safety-classifier_gte-large-en-v1.5. Schema Column Type Description id string Row identifier (matches source dataset) safety_score int8 Argmax safety class (0-5) safety_probs list[float32] Full 6-class probability distribution Safety scale Score Label Count Percentage 0 safe 302,972,734 77.39% 1… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/safety-classifications.texttext-classification100M<n<1B1 likes801 downloads2mo agoHugging Face21ourafla /Mental-Health_Text-Classification_Dataset Mental Health Text Classification Dataset (4-Class) Dataset Description This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models. The repository includes: An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.texttext-classification10K<n<100K13 likes777 downloads10mo agoHugging Face22mesolitica /Zeroshot-Audio-Classification-Instructions Zeroshot-Audio-Classification-Instructions Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label, VGGSound FSD50k Nonspeech7k urbansound8K VocalSound Emotion Gender ESD Emotion Age Language TAU Urban Acoustic Scenes 2022 CochlScene BirdCLEF_2021 EmoBox AudioSet We also converted huge WAV files into MP3 16k sample rate to reduce storage size.To prevent leakage, please do not include test set in training session.… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.audio1M<n<10M7 likes743 downloads1y agoHugging Face23ibm-nasa-geospatial /multi-temporal-crop-classification Dataset Card for Multi-Temporal Crop Classification Dataset Summary This dataset contains temporal Harmonized Landsat-Sentinel imagery of diverse land cover and crop type classes across the Contiguous United States for the year 2022. The target labels are derived from USDA's Crop Data Layer (CDL). It's primary purpose is for training segmentation geospatial machine learning models. Dataset Structure TIFF Files Each tiff file covers a… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/multi-temporal-crop-classification.text1K<n<10K29 likes741 downloads2y agoHugging Face24ClaudiaRichard /mbti_classification_dataset_fullPosts MBTI Classification Dataset (Full Posts) A dataset of 8,675 personality-forum posts labeled with Myers-Briggs Type Indicator (MBTI) dichotomies, built for training and evaluating text-based personality classification models. Each row contains one user's concatenated forum posts plus binary labels for all four MBTI dimensions. Dataset Structure Splits: train (5,205 rows), test (2,082 rows), validation (1,388 rows) Fields: Field Type Description I/E int64… See the full description on the dataset page: https://huggingface.co/datasets/ClaudiaRichard/mbti_classification_dataset_fullPosts.tabulartext-classification1K<n<10K1 likes740 downloads1d agoHugging Face25keremberke /chest-xray-classification Dataset Labels ['NORMAL', 'PNEUMONIA'] Number of Images {'train': 4077, 'test': 582, 'valid': 1165} How to Use Install datasets: pip install datasets Load the dataset: from datasets import load_dataset ds = load_dataset("keremberke/chest-xray-classification", name="full") example = ds['train'][0] Roboflow Dataset Page https://universe.roboflow.com/mohamed-traore-2ekkp/chest-x-rays-qjmia/dataset/2 Citation… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/chest-xray-classification.imageimage-classification1K<n<10K28 likes673 downloads4y agoHugging Face26iitolstykh /LLMTrace_classification LLMTrace - Classification Dataset 🌐 LLMTrace Website | 📜 LLMTrace Paper on arXiv | 🤗 LLMTrace - Detection Dataset | 🤗 GigaCheck classification model | This repository contains the Classification portion of the LLMTrace project. This dataset is specifically designed for the binary classification of texts as either human-written or AI-generated. For full details on the data collection methodology, statistics, and experiments, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/iitolstykh/LLMTrace_classification.text100K<n<1M2 likes642 downloads10mo agoHugging Face27mteb /Vehicle_sounds_classification_datasetaudio1K<n<10K1 likes632 downloads8mo agoHugging Face28kenhktsui /code-natural-language-classification-datasetSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'] + sampling from minipile. It is intended to be used for training code natural language classifier. texttext-classification1M<n<10M0 likes594 downloads2y agoHugging Face29dolphinteam /OpenWhistle-Classification-Finetuning OpenWhistle Classification Finetuning Dataset dolphinteam/OpenWhistle-Classification-Finetuning is the public classification finetuning dataset used for dolphin whistle identity classification. It contains short whistle clips, whistle-level metadata, fundamental-frequency tracks, rendered F0 spectrograms, and integer class labels. The main reviewer-facing subset is the balanced balanced config. It contains six classes: NSW_1 (label=0) SW_Luna (label=1) SW_Nana (label=2) SW_Neo… See the full description on the dataset page: https://huggingface.co/datasets/dolphinteam/OpenWhistle-Classification-Finetuning.audioaudio-classification10K<n<100K2 likes576 downloads10d agoHugging Face30limsc /fr-nfr-classificationtextn<1K2 likes565 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.