Team Ai
Datasetpublic

AmazonScience/massive

MASSIVE is a parallel dataset of > 1M utterances across 51 languages with annotations for the Natural Language Understanding tasks of intent prediction and slot annotation. Utterances span 60 intents and include 55 slot types. MASSIVE was created by localizing the SLURP dataset, composed of general Intelligent Voice Assistant single-shot interactions.

sourceHugging Facecc-by-4.0updated 4y agoView on Hugging Face
101likes16kdownloads
CITATION.cff50 linesDownload Raw Back to root
1cff-version: 1.0.02message: "MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages"3authors:4- family-names: "FitzGerald"5 given-names: "Jack"6- family-names: "Hench"7 given-names: "Christopher"8- family-names: "Peris"9 given-names: "Charith"10- family-names: "Mackie"11 given-names: "Scott"12- family-names: "Rottmann"13 given-names: "Kay"14- family-names: "Sanchez"15 given-names: "Ana"16- family-names: "Nash"17 given-names: "Aaron"18- family-names: "Urbach"19 given-names: "Liam"20- family-names: "Kakarala"21 given-names: "Vishesh"22- family-names: "Singh"23 given-names: "Richa"24- family-names: "Ranganath"25 given-names: "Swetha"26- family-names: "Crist"27 given-names: "Laurie"28- family-names: "Britan"29 given-names: "Misha"30- family-names: "Leeuwis"31 given-names: "Wouter"32- family-names: "Tur"33 given-names: "Gokhan"34- family-names: "Natarajan"35 given-names: "Prem"36title: "MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages"37version: 1.0.038date-released: 2022-04-1839url: "https://arxiv.org/abs/2204.08582"40 41  42  43 44  45  46 47  48  49 50