Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /ade_corpus_v2_classification ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data. This is a dataset for classification if a sentence is ADE-related (True) or not (False). Train size: 17,637 Test size: 5,879 Source dataset Paper text10K<n<100K6 likes2.4k downloads4y agoHugging Face02iitolstykh /LLMTrace_classification LLMTrace - Classification Dataset 🌐 LLMTrace Website | 📜 LLMTrace Paper on arXiv | 🤗 LLMTrace - Detection Dataset | 🤗 GigaCheck classification model | This repository contains the Classification portion of the LLMTrace project. This dataset is specifically designed for the binary classification of texts as either human-written or AI-generated. For full details on the data collection methodology, statistics, and experiments, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/iitolstykh/LLMTrace_classification.text100K<n<1M2 likes650 downloads10mo agoHugging Face03ai-forever /ru-reviews-classificationtexttext-classification10K<n<100K6 likes535 downloads2y agoHugging Face04ai-forever /kinopoisk-sentiment-classificationtexttext-classification10K<n<100K7 likes468 downloads2y agoHugging Face05ai-forever /headline-classificationtexttext-classification10K<n<100K1 likes409 downloads2y agoHugging Face06ai-forever /ru-scibench-grnti-classificationtexttext-classification10K<n<100K0 likes407 downloads2y agoHugging Face07hasankursun /multilingual-safety-classification-dataset Multilingual Safety Classification Dataset A multilingual dataset for safety classification across 60 languages, created by Hasan Kurşun through machine translation of English safety prompts using NLLB-200-3.3B. Dataset Details Processed by: Hasan KurşunAuthor: Hasan KurşunYear: 2025Source Dataset: mvrcii/safety-moderation-benchmarkTranslation Model: facebook/nllb-200-3.3B Languages (60) African Languages (16): Amharic, Hausa, Kinyarwanda, Luganda… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/multilingual-safety-classification-dataset.texttext-classification100K<n<1M4 likes363 downloads4mo agoHugging Face08agentlans /en-document-classification English Document Classification Dataset This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. Dataset Summary The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.texttext-classification1M<n<10M1 likes346 downloads1mo agoHugging Face09ai-forever /ru-scibench-oecd-classificationtexttext-classification10K<n<100K0 likes323 downloads2y agoHugging Face10samscript18 /adaption-defi-wallet-risk-classification This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-defi_wallet_risk_classification This dataset contains prompt-completion pairs for classifying the 14-day risk outcomes of DeFi wallets on various EVM networks. Each entry provides behavioral features such as transaction counts, action ratios, and concentration metrics within a specific feature window to predict a binary risk label. The completions offer a concise justification for… See the full description on the dataset page: https://huggingface.co/datasets/samscript18/adaption-defi-wallet-risk-classification.textn<1K0 likes276 downloads3mo agoHugging Face11ai-forever /inappropriateness-classificationtexttext-classification10K<n<100K1 likes256 downloads2y agoHugging Face12gehits /Chinese-Legal-Case-Classification-Datasettext10K<n<100K9 likes249 downloads9mo agoHugging Face13asahi417 /multi-domain-document-classification multi_domain_document_classification Multi-domain document classification datasets. Biomedical: chemprot, rct-sample Computer Science: citation_intent, sciie Customer Review: amcd, yelp_review Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train. chemprot citation_intent hyperpartisan_news rct_sample sciie amcd yelp_review tweet_eval_irony tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.text10K<n<100K0 likes232 downloads4y agoHugging Face14helenqu /astro-classification-redshifts AstroClassification and Redshifts Datasets This dataset was used for the AstroClassification and Redshifts introduced in Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations. This is a dataset of simulated astronomical time-series (e.g., supernovae, active galactic nuclei), and the task is to classify the object type (AstroClassification) or predict the object's redshift (Redshifts). Repository: https://github.com/helenqu/connect-later Paper: will be… See the full description on the dataset page: https://huggingface.co/datasets/helenqu/astro-classification-redshifts.tabular100K<n<1M1 likes204 downloads3y agoHugging Face15vhands /audio-event-classification-post-public audio-event-classification-post-public Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.textaudio-classification100K<n<1M1 likes204 downloads3mo agoHugging Face16ai-forever /sensitive-topics-classificationtexttext-classification10K<n<100K1 likes189 downloads2y agoHugging Face17semaj83 /ctmatch_classificationCTMatch Classification Dataset This is a combined set of 2 labelled datasets of: topic (patient descriptions), doc (clinical trials documents - selected fields), and label ({0, 1, 2}) triples, in jsonl format. (Somewhat of a duplication of some of the ir_dataset also available on HF.) These have been processed using ctproc, and in this state can be used by various tokenizers for fine-tuning (see ctmatch for examples). These 2 datasets contain no patient identifying information are openly… See the full description on the dataset page: https://huggingface.co/datasets/semaj83/ctmatch_classification.texttext-classification10K<n<100K2 likes186 downloads3y agoHugging Face18scikit-fingerprints /TDC_b3db_classificationn<1K0 likes181 downloads2y agoHugging Face19agentlans /en-document-topic-classification English Document Topic Classification Dataset English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics. Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks. Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation). Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.text1M<n<10M0 likes171 downloads1mo agoHugging Face20agentlans /en-document-format-classification English Document Format Classification Dataset English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats. Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks. Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage). Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.text100K<n<1M0 likes151 downloads1mo agoHugging Face21MCINext /pershop-classification@article{mahmoudi2024pershop, title={PerSHOP--A Persian dataset for shopping dialogue systems modeling}, author={Mahmoudi, Keyvan and Faili, Heshaam}, journal={arXiv preprint arXiv:2401.00811}, year={2024} } text10K<n<100K0 likes141 downloads1y agoHugging Face22bench-labs /slop-classification Slop classifier dataset A human-annotated dataset for studying and classifying AI-generated text that people perceive as “AI slop.” The dataset is built from samples collected from existing public datasets and annotated through the Bench Labs SlopFinder interface. Slop score Each sample receives a score based on human votes: -1 = definitely slop 0 = undecided / neutral +1 = not slop at all The score represents human judgment, not an objective measure of quality… See the full description on the dataset page: https://huggingface.co/datasets/bench-labs/slop-classification.tabularn<1K11 likes129 downloads1d agoHugging Face23marcelsun /wos_hierarchical_multi_label_text_classificationIntroduced by du Toit and Dunaiski (2024) Introducing Three New Benchmark Datasets for Hierarchical Text Classification. The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created… See the full description on the dataset page: https://huggingface.co/datasets/marcelsun/wos_hierarchical_multi_label_text_classification.texttext-classification100K<n<1M0 likes114 downloads2y agoHugging Face24fahmidiqbal /tlink-classificationtext10K<n<100K0 likes112 downloads5mo agoHugging Face25MCINext /style-classificationtext1K<n<10K0 likes108 downloads1y agoHugging Face26MCINext /synthetic-persian-text-keyword-pair-classification Dataset Summary Synthetic Persian Text-Keywords Pair Classification (SynPerTextKeywordsPC) is a Persian (Farsi) dataset developed for the Pair Classification task. The dataset focuses on determining whether a keyword or short phrase is relevant to a longer Persian text passage. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically created using GPT-4o-mini. Language(s): Persian (Farsi) Task(s): Pair Classification (Text–Keyword Relevance)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-text-keyword-pair-classification.text10K<n<100K0 likes101 downloads1y agoHugging Face27imnim /multiclass-email-classificationThis dataset comprises of more than 2000 emails across multiple categories, which can he helpful for tasks like LLM training and fine-tuning. The dataset is also provided with a python script that would generate emails automatically The dataset contains email across 10 different categories namely, "Business", "Personal", "Promotions", "Customer Support", "Job Application", "Finance & Bills", "Events & Invitations", "Travel & Bookings", "Reminders", "Newsletters" Total emails: 2105 Label… See the full description on the dataset page: https://huggingface.co/datasets/imnim/multiclass-email-classification.text1K<n<10K4 likes93 downloads1y agoHugging Face28CJJones /Wikipedia_RAG_QA_Classification 🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training 📊 Dataset Description This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning. 🖥️ Demo Interface: Discord Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.tabulartext-generation100K<n<1M1 likes89 downloads7mo agoHugging Face29MCINext /synthetic-persian-chatbot-rag-faq-pair-classification Dataset Summary Synthetic Persian Chatbot RAG FAQ Pair Classification (SynPerChatbotRAGFAQPC) is a Persian (Farsi) dataset for the Pair Classification task, specifically designed for evaluating Retrieval-Augmented Generation (RAG) chatbot systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini. The dataset measures a model’s ability to assess whether a given FAQ (question–answer pair) is relevant to a new user… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-faq-pair-classification.text1K<n<10K0 likes88 downloads1y agoHugging Face30referencesource /hazardous-area-substance-classification Gas group, temperature class and autoignition temperature by substance Canonical, always-current version: https://referencesource.org/hazardous-area-substance-classification/ Machine-readable: https://referencesource.org/hazardous-area-substance-classification/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-06 Stale after: 2028-08-05 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 73 Hazardous… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/hazardous-area-substance-classification.textn<1K0 likes88 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.