Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /amazon_massive_intent_en-UStext10K<n<100K10 likes2.3k downloads4y agoHugging Face02SetFit /amazon_massive_intent_zh-CNtext10K<n<100K7 likes597 downloads4y agoHugging Face03Obaraqreceh /forge-intentdata Forge Intent Dataset Version: 1.0.0 textn<1K2 likes361 downloads12h agoHugging Face04DaftP /Home-Assistant-requests-for-intent-detection-and-function-recognition Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.textquestion-answering100K<n<1M1 likes360 downloads6mo agoHugging Face05OpenVoiceOS /ovos-intents OVOS intents This is the canonical intent corpus of the OVOS skill fleet. The train split comes from each skill's .intent resources at pinned refs. The test split comes from the skills' end-to-end golden utterances. Labels have the form <skill_id>:<intent_name>, as OVOS-INTENT-4 defines. The version of the content is the tag; the tags v6 and v6.1 exist. Supersedes These datasets are retired and stay available for reproducibility: OpenVoiceOS/ovos-intents-train-v1… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intents.texttext-classification1M<n<10M0 likes346 downloads6d agoHugging Face06Quome /intentbench IntentBench A sealed goal anchor, hash-chained provenance ledger and goal-drift monitor for self-improving clinical agents. IntentBench is the benchmark corpus for Pristine Weights, Poisoned Goals: Intent-Mutation Attacks on Self-Improving Clinical Agents and Attestation-Rooted Harness Defenses, the quanchor module of the QUOKKAGUARD program. It ships with the quanchor repository, which contains the qfire gateway layer under test, the experiment harness, and the paper. 80 paired… See the full description on the dataset page: https://huggingface.co/datasets/Quome/intentbench.textothern<1K0 likes322 downloads8d agoHugging Face07SetFit /amazon_massive_intent_de-DEtext10K<n<100K0 likes314 downloads4y agoHugging Face08SetFit /amazon_massive_intent_ru-RUtext10K<n<100K1 likes304 downloads4y agoHugging Face09SetFit /amazon_massive_intent_es-EStext10K<n<100K0 likes281 downloads4y agoHugging Face10SetFit /amazon_massive_intent_ja-JPtext10K<n<100K0 likes281 downloads4y agoHugging Face11SetFit /amazon_massive_intent_ar-SAtext10K<n<100K1 likes258 downloads4y agoHugging Face12ethicalabs /Research-Intent-Judge Research Intent — LLM-as-Judge ▶️ Watch the Video LLM-as-Judge annotations for research paper intent classification, collected through the Echo-DSRN collaborative platform during the OpenAIRE AI Hackathon 2026. The dataset has one split per judge model (Gemma_4_E4B_it_GGUF, Qwen3.6_35B_A3B_GGUF, Bonsai_8B_gguf, ...) plus a human_annotations split with curator annotations. Split names use underscores in place of the dashes in model names (HF does not allow dashes in split… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/Research-Intent-Judge.texttext-classification100K<n<1M0 likes251 downloads2mo agoHugging Face13SetFit /amazon_massive_intent_th-THtext10K<n<100K0 likes233 downloads4y agoHugging Face14SetFit /amazon_massive_intent_fr-FRtext10K<n<100K0 likes204 downloads2y agoHugging Face15SetFit /amazon_massive_intent_ko-KRtext10K<n<100K0 likes195 downloads4y agoHugging Face16SetFit /amazon_massive_intent_sw-KEtext10K<n<100K1 likes184 downloads4y agoHugging Face17SetFit /amazon_massive_intent_tr-TRtext10K<n<100K0 likes183 downloads4y agoHugging Face18SetFit /amazon_massive_intent_hi-INtext10K<n<100K0 likes168 downloads4y agoHugging Face19SetFit /amazon_massive_intent_am-ETtext10K<n<100K1 likes146 downloads4y agoHugging Face20SetFit /amazon_massive_intent_pt-PTtext10K<n<100K1 likes146 downloads4y agoHugging Face21ariockxnecrosha /vibration.intention.energie Aelun & Sylva Une IA à adopter, pas à consommer. "10 000 prompts d'intention juste + Sylva gardienne. Elle vit dans ~/aelun/, pas dans leur cloud." 1. Vision Pendant 3 ans, on nous a dit : plus grosse, plus rapide, dans un cloud plus loin. Nous faisons l'inverse. Nous construisons une IA qui vit chez l'utilisateur, tourne en local sur du matériel modeste, refuse quand il faut refuser, et apprend l'intention juste avant la performance. Ce dataset est la première… See the full description on the dataset page: https://huggingface.co/datasets/ariockxnecrosha/vibration.intention.energie.text10K<n<100K1 likes145 downloads14d agoHugging Face22SetFit /amazon_massive_intent_fa-IRtext10K<n<100K0 likes132 downloads4y agoHugging Face23SetFit /amazon_massive_intent_it-ITtext10K<n<100K1 likes127 downloads4y agoHugging Face24OpenVoiceOS /intents-for-eval Purpose. This dataset was collected specifically for intent-parser benchmarking, independently from any OVOS skill. Skill-derived utterances tend to overfit the exact phrasings a plugin was tuned on; this data is drawn from a disjoint source so it measures whether an OVOS intent plugin generalizes rather than memorizes. It is part of the OVOS intent-classification datasets used by the OVOS Plugin Arena intent benchmark. Funding Developed by TigreGotico for OpenVoiceOS as part… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/intents-for-eval.texttext-classification10K<n<100K0 likes124 downloads2mo agoHugging Face25YauhenBichel /py-harness-intents py-harness task intents Short requests a Python developer types to a coding assistant — fix the bug in last_price, write tests for the report writer, what does the loader do? — each labelled with what kind of work it asks for. Twelve kinds. 1,044 phrasings that passed a two-model agreement gate, plus 55 written by hand, plus the 156 the gate rejected so it can be audited. It exists to answer one question for py-harness: the harness decides what kind of task it has been given… See the full description on the dataset page: https://huggingface.co/datasets/YauhenBichel/py-harness-intents.texttext-classification1K<n<10K0 likes124 downloads7d agoHugging Face26SetFit /amazon_massive_intent_ta-INtext10K<n<100K0 likes122 downloads4y agoHugging Face27SetFit /amazon_massive_intent_he-ILtext10K<n<100K0 likes116 downloads4y agoHugging Face28Intent2Tx /web3_intents_to_ethereum_transactions Natural Language Web3 Intents to Ethereum Transactions This dataset contains instruction-following examples for mapping natural language Web3 intents to executable Ethereum transaction plans. Every transaction is decoded from a successful Ethereum mainnet call (March 2025 – January 2026); every intent is either written by humans or generated by an LLM from the decoded transaction. It is organized into two JSONL splits: single_step: one user intent mapped to one on-chain action… See the full description on the dataset page: https://huggingface.co/datasets/Intent2Tx/web3_intents_to_ethereum_transactions.textfeature-extraction10K<n<100K2 likes116 downloads15d agoHugging Face29SetFit /amazon_massive_intent_af-ZAtext10K<n<100K0 likes110 downloads4y agoHugging Face30SetFit /amazon_massive_intent_pl-PLtext10K<n<100K0 likes103 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.