datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amazon_massive_intent
MassiveIntentClassification
An MTEB dataset
Massive Text Embedding Benchmark
MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages
Task category
t2c
Domains
Spoken
Reference
https://arxiv.org/abs/2204.08582
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MassiveIntentClassification"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/amazon_massive_intent.amazon_massive_intent_en-USbuy_sell_intentmtop_intent
MTOPIntentClassification
An MTEB dataset
Massive Text Embedding Benchmark
MTOP: Multilingual Task-Oriented Semantic Parsing
Task category
t2c
Domains
Spoken, Spoken
Reference
https://arxiv.org/pdf/2008.09335.pdf
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["MTOPIntentClassification"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/mtop_intent.snips_built_in_intents
Dataset Card for Snips Built In Intents
Dataset Summary
Snips' built in intents dataset was initially used to compare different voice assistants and released as a public dataset hosted at
https://github.com/sonos/nlu-benchmark in folder 2016-12-built-in-intents. The dataset contains 328 utterances over 10 intent classes.
A related Medium post is https://medium.com/snips-ai/benchmarking-natural-language-understanding-systems-d35be6ce568d.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sonos-nlu-benchmark/snips_built_in_intents.IntentionDatasetMT-intents-dataset-pt-PT
Retired. This dataset is superseded by OpenVoiceOS/ovos-intents. It stays available for reproducibility and receives no updates.
amazon_massive_intent_zh-CNIntentQAMovens-Intent Movens-Intent
An omni-modal benchmark for evaluating intent-to-humanoid motion generation
Evaluation-only split. Movens-Intent contains fixed benchmark samples selected from the Movens training-data pool. It is intended to evaluate models that were not trained on these exact samples. Any overlap with a model's training data must be disclosed.
Overview
Movens-Intent evaluates whether a humanoid motion model can follow intent expressed through four input… See the full description on the dataset page: https://huggingface.co/datasets/wendell0218/Movens-Intent.IntentGrasp
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
Paper: https://arxiv.org/abs/2605.06832
Authors: Yuwei Yin, Chuyuan Li, Giuseppe Carenini
Institute: UBC NLP Group, Department of Computer Science, University of British Columbia
Keywords: Intent Understanding, Dataset, Benchmark, LLM, Evaluation, Intentional Fine-Tuning
Abstract:
Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language… See the full description on the dataset page: https://huggingface.co/datasets/yuweiyin/IntentGrasp.transfer-bench-intents-v2-packed
transfer-bench-intents-v2
CUDA→Ascend NPU 迁移评测任务的 intent 清单数据集(第二代):3619 个真实开源 repo,
每个含 5–25 条经校验的 tier1/tier2 测试意图(intents.json)+ analyst 分析底稿(notes.md),
以及迁移任务元数据与打包的原始 repo 快照。
选集:rosetta-selector(transfer-bench-Rosetta/rosetta-selector/)fastpath
确定性筛选管线产出——口径"PyTorch 项目且未适配 NPU",零 agent 成本;
候选集 rosetta-selector/runs/20260831-154337/candidates
生成:transfer-bench-Rosetta/intents-generation/gen_intents.py,模型 kimi/kimi-k2.7-code,
source run… See the full description on the dataset page: https://huggingface.co/datasets/foreverCuSO4/transfer-bench-intents-v2-packed.Uddessho-Bangla-Multimodal-Intent-Classification
📊 Uddessho Dataset — Multimodal Author Intent Classification
Uddessho (meaning "Intent" in English) is a multimodal dataset created for author intent classification in the low-resource Bangla language.It contains 3,048 social media posts (text + images) labeled into six distinct intent types.
🏷️ Intent Categories & Label Mapping
Label ID
Class Name
0
Advocative
1
Controversial
2
Exhibitionist
3
Expressive
4
Informative
5
Promotive
📂… See the full description on the dataset page: https://huggingface.co/datasets/Mukaffi28/Uddessho-Bangla-Multimodal-Intent-Classification.turkish-intent-classification-1m
Turkish Intent Classification 1M v2
Yirmi destek niyetini kapsayan slot çeşitlendirmeli Türkçe sınıflandırma verisi.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, text, label
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-intent-classification-1m.forge-intentdata
Forge Intent Dataset
Version: 1.0.0
Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.ovos-intents
OVOS intents
This is the canonical intent corpus of the OVOS skill fleet. The train split comes from each
skill's .intent resources at pinned refs. The test split comes from the skills' end-to-end
golden utterances. Labels have the form <skill_id>:<intent_name>, as OVOS-INTENT-4 defines.
The version of the content is the tag; the tags v6 and v6.1 exist.
Supersedes
These datasets are retired and stay available for reproducibility:
OpenVoiceOS/ovos-intents-train-v1… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intents.online-shoppers-intentionintentbench
IntentBench
A sealed goal anchor, hash-chained provenance ledger and goal-drift monitor for self-improving clinical agents.
IntentBench is the benchmark corpus for Pristine Weights, Poisoned Goals: Intent-Mutation Attacks on Self-Improving Clinical Agents and Attestation-Rooted Harness Defenses, the quanchor module of the QUOKKAGUARD program. It ships with the quanchor repository, which contains the qfire gateway layer under test, the experiment harness, and the paper.
80 paired… See the full description on the dataset page: https://huggingface.co/datasets/Quome/intentbench.peak-intent-50amazon_massive_intent_de-DEIntentEmotionamazon_massive_intent_ru-RUamazon_massive_intent_es-ESamazon_massive_intent_ja-JPamazon_massive_intent_ar-SAIntentBench
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
Paper: HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
Code: GitHub Repository
Hugging Face Dataset: IntentBench (a key benchmark introduced with HumanOmniV2)
Other Resource: ModelScope
👀 HumanOmniV2 Overview
With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability… See the full description on the dataset page: https://huggingface.co/datasets/PhilipC/IntentBench.Research-Intent-Judge
Research Intent — LLM-as-Judge
▶️ Watch the Video
LLM-as-Judge annotations for research paper intent classification, collected
through the Echo-DSRN collaborative platform during the OpenAIRE AI Hackathon 2026.
The dataset has one split per judge model (Gemma_4_E4B_it_GGUF,
Qwen3.6_35B_A3B_GGUF, Bonsai_8B_gguf, ...) plus a human_annotations
split with curator annotations. Split names use underscores in place of the
dashes in model names (HF does not allow dashes in split… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/Research-Intent-Judge.ovos-tts-bench-intents-for-eval-prompts
OVOS tts bench — intents-for-eval-prompts
Synthesised clips (one per prompt) predictions of the registered
OVOS Plugin Arena
tts fighters over
OpenVoiceOS/intents-for-eval.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-intents-for-eval-prompts.amazon_massive_intent_th-TH
