AmazonScience/massive
MASSIVE is a parallel dataset of > 1M utterances across 51 languages with annotations for the Natural Language Understanding tasks of intent prediction and slot annotation. Utterances span 60 intents and include 55 slot types. MASSIVE was created by localizing the SLURP dataset, composed of general Intelligent Voice Assistant single-shot interactions.
10116k
1from datasets import load_dataset2 3source = "AmazonScience/massive"4 5#Individual locales are downloaded from version 1.1 which has Catalan (ca-ES) added in6dataset = load_dataset(source, "en-US", download_mode="force_redownload")7print(dataset)8print(dataset["train"][0])9 10#Using version 1.1 which has Catalan (ca-ES) added in11dataset = load_dataset(source, "all_v1.1", download_mode="force_redownload")12print(dataset)13print(dataset["train"][0])14 15#Using version 1.016dataset = load_dataset(source, "all", download_mode="force_redownload")17print(dataset)18print(dataset["train"][0])19 