datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Uddessho-Bangla-Multimodal-Intent-Classification
📊 Uddessho Dataset — Multimodal Author Intent Classification
Uddessho (meaning "Intent" in English) is a multimodal dataset created for author intent classification in the low-resource Bangla language.It contains 3,048 social media posts (text + images) labeled into six distinct intent types.
🏷️ Intent Categories & Label Mapping
Label ID
Class Name
0
Advocative
1
Controversial
2
Exhibitionist
3
Expressive
4
Informative
5
Promotive
📂… See the full description on the dataset page: https://huggingface.co/datasets/Mukaffi28/Uddessho-Bangla-Multimodal-Intent-Classification.turkish-intent-classification-1m
Turkish Intent Classification 1M v2
Yirmi destek niyetini kapsayan slot çeşitlendirmeli Türkçe sınıflandırma verisi.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, text, label
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-intent-classification-1m.IntentClassification_Dataset_for_AI_Assistant_Prompt_Routing_HindiChildren_Intent_Classification
MAMA Communicative Intent Dataset (INCA-A Annotated)
Overview
The MAMA Communicative Intent Dataset is a linguistically annotated corpus of child utterances designed to support research in child-centred Natural Language Processing (NLP) and communicative intent recognition in early language development.
The dataset contains 10,800 child utterances annotated using the INCA Communicative Coding System (Ninio et al., 1994), a developmental framework that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/Wajinimi/Children_Intent_Classification.intent-classification-en-frmtop_intent_classificationThis dataset contains annotated utterances from 6 languages, including Thai,
for semantic parsing. Queries corresponding to the chosen domains are crowdsourced.
Two subsets are included in this dataset: 'domain' (eg. 'news', 'people', 'weather')
and 'intent' (eg. 'GET_MESSAGE', 'STOP_MUSIC', 'END_CALL')amazon_massive_intent_fr_prompt_intent_classification
amazon_massive_intent_fr_prompt_intent_classification
Summary
amazon_massive_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 555,000 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset amazon_massive_intent_fr-FR by FitzGerald et al..
A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_massive_intent_fr_prompt_intent_classification.massive-intent-ind-classificationref: https://huggingface.co/datasets/mteb/amazon_massive_intent
Email_Intent_Classification
Dataset Information
This is a dataset of English sentences used in emails with six basic categories: request, informational, transaction, feedback.
An example looks as follows: {"Email": "Your subscription renewal is confirmed. Thank you for staying with us!", "Intent": "Transaction"}
Dataset Sources
Instances generated and annotated by ChatGPT 4.
Uses
Demo for email intent classification tasks.
banking-intent-classification
Banking Intent Classification Dataset
This dataset contains text samples for banking intent classification tasks.
Dataset Description
The dataset consists of customer queries/messages related to banking services, each labeled with an intent category.
Usage
import pandas as pd
# Load the dataset
df = pd.read_csv("hf://datasets/Cleanlab/banking-intent-classification/banking-intent-classification.csv")
print(df.head())
License
MIT License
intent_classificationatis_intent_classification_translatedIntent-Classification-large
Dataset Card for "Intent-Classification-large"
More Information needed
mtop_domain_intent_fr_prompt_intent_classification
mtop_domain_intent_fr_prompt_intent_classification
Summary
mtop_domain_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 497,100 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset mtop_domain Haoran Li et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/mtop_domain_intent_fr_prompt_intent_classification.massive-intent-zsm-classificationmassive-intent-vie-classificationmassive-intent-ind-classificationmassive-intent-zsm-classification
Dataset Card for "ms-intent-classification"
More Information needed
massive-intent-tha-classificationbitext-customer-support-intent-classificationmassive-intent-fil-classificationmassive-intent-fil-classification
MassiveIntent_fil_Classification
Deduplicated copy of kornwtp/massive-intent-fil-classification.
Splits
split
rows
test
2,943
train
11,173
validation
2,014
massive-intent-khm-classificationmassive-intent-tam-classificationmassive-intent-vie-classification
MassiveIntent_vie_Classification
Deduplicated copy of kornwtp/massive-intent-vie-classification.
Splits
split
rows
test
2,935
train
11,126
validation
2,020
intent-classification-60k
Dataset Card for Intent Classification Dataset
Provide a quick summary of the dataset.
This dataset contains 60,000 unique user prompts classified into six distinct intent categories using GLM-5-Turbo, a state-of-the-art large language model.
Dataset Details
Dataset Description
A curated dataset of 60,000 user prompts labeled with intent categories: CODING, CHAT, REASONING, SIMPLE, TOOL, and BASIC. Each prompt has been consistently labeled using GLM-5-Turbo… See the full description on the dataset page: https://huggingface.co/datasets/atekrugis/intent-classification-60k.med_intent_classificationmassive-intent-tam-classification
Dataset Card for "ta-intent-classification"
More Information needed
intent-classification-v9-boundary-baseline
Dataset Card: Intent Classification Baseline v9 Boundary
Dataset Summary
This dataset is the baseline training pool used for the v9 boundary-focused fastText router.
File: final_dataset_augmented_v9_boundary.csv
Rows: 65,595
Columns:
prompt (string)
final_label (string)
Labels:
BASIC
SIMPLE
CHAT
REASONING
TOOL
CODING
Dataset Composition
Label distribution in this baseline CSV:
CODING: 23,331
CHAT: 16,037
REASONING: 11,972
SIMPLE: 9,939
TOOL: 3,333… See the full description on the dataset page: https://huggingface.co/datasets/atekrugis/intent-classification-v9-boundary-baseline.email-intent-classification
