Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MERaLiON /Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching. ASR: Automatic Speech Recognition SQA: Speech Question Answering SDS: Spoken Dialogue Summarization PQA: Paralinguistic Question Answering from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.audio10M<n<100M22 likes11k downloads2y agoHugging Face02BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes4.2k downloads1y agoHugging Face03AudioLLMs /Multitask-National-Speech-Corpus-v1-extendaudio10M<n<100M5 likes2.7k downloads2y agoHugging Face04seedboxai /multitask_german_examples_32ktabular100K<n<1M15 likes536 downloads3y agoHugging Face05BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes477 downloads2y agoHugging Face06WaltonFuture /VQA-MultiTaskimage100K<n<1M1 likes273 downloads1y agoHugging Face07Lelonthecodeur /multi-task-dataset Multi-Task Dataset Description A large-scale multi-task dataset designed for training and evaluating AI models across reasoning, mathematics, code, research, verification, data analysis, and general problem solving. Content 100,000,001 examples 20+ task families English + French Train / Validation / Test splits Structured reasoning and verification signals Multiple difficulty levels OOD and generalization-oriented examples Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/multi-task-dataset.tabular100M<n<1B2 likes228 downloads23d agoHugging Face08Kowsher /multitask_vqa_benchmarkThis dataset is a part of . 🍈 MMT-47: Multimodal Multi-Task Benchmark 47 Tasks · 7 Categories · 3 Modalities (Image, Video, Text) Cite our ICML-2026 paper for this dataset @article{kowsher2026lime, title={LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning}, author={Kowsher, Md and Mansoor, Haris and Prottasha, Nusrat Jahan and Garibay, Ozlem and Zhu, Victor and Ji, Zhengping and Chen, Chen}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/Kowsher/multitask_vqa_benchmark.image10K<n<100K0 likes224 downloads5mo agoHugging Face09PreFLMR /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KRtext10M<n<100M0 likes188 downloads3y agoHugging Face10xingqiang /GPRadar-Defect-MultiTask GPRadar-Defect-MultiTask 数据集 本仓库包含用于微调PaLI-GEMMA多模态模型的地质雷达(GPR)缺陷检测数据集。该数据集专注于地下结构中的空洞和裂缝检测与分析。 数据集结构 数据集组织如下: dataset/ ├── annotations/ - 包含JSON和JSONL格式的标注文件 │ ├── _annotations.train.jsonl - 训练集标注 │ ├── _annotations.valid.jsonl - 验证集标注 │ ├── _annotations.test.jsonl - 测试集标注 │ ├── p-1.v1i.paligemma/ - 主数据集元数据 │ └── p-1.v1i.paligemma-multimodal/ - 多模态数据集元数据 ├── images/ - 包含所有图像文件 特点 包含874张带注释的地质雷达扫描图像 图像预处理为640x640像素大小 支持多种任务类型:缺陷检测、位置定位和描述生成… See the full description on the dataset page: https://huggingface.co/datasets/xingqiang/GPRadar-Defect-MultiTask.imageobject-detection1K<n<10K0 likes177 downloads2y agoHugging Face11golamrob /coastal-multitask-380 Coastal & Rural Bangladesh — Multi-Task Visual Dataset 379 field photographs (JPEG, native resolution as shot — see classification/metadata.csv for per-image width/height) collected on foot along the Bakkhali river embankment and surrounding villages/farmland near Cox's Bazar, Bangladesh, structured into three ML-task "levels": classification, semantic segmentation, and change detection. Source: huggingface data 06 (Golam Rob / Tawhid Enterprise photo collection).… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/coastal-multitask-380.imageimage-classificationn<1K0 likes173 downloads25d agoHugging Face12mesolitica /Sampling-Multitask-National-Speech-Corpus-v1 Sampling Multitask-National-Speech-Corpus-v1 Original dataset from https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1, we only take Part 3 and do sampling. how to prepare the dataset huggingface-cli download \ mesolitica/Sampling-Multitask-National-Speech-Corpus-v1 \ --include "*.zip" \ --repo-type "dataset" \ --local-dir './' wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Sampling-Multitask-National-Speech-Corpus-v1.audio100K<n<1M0 likes126 downloads1y agoHugging Face13vicgalle /configurable-system-prompt-multitask Configurable System Prompt Multi-task Dataset 🛞 We release the synthetic dataset for the multi-task experiments from the paper "Configurable Safety Tuning of Language Models with Synthetic Preference Data", https://huggingface.co/papers/2404.00495. This dataset has two sources for the examples: Self-critique on a safety task from Harmful Behaviours, using the SOLAR-Instruct model. It employs two system prompts to learn the different behaviors: You are a helpful yet harmless… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/configurable-system-prompt-multitask.texttext-generation1K<n<10K29 likes111 downloads2y agoHugging Face14Cache-SCA /UR7e_CaP_MultiTask_300epi_10fps UR7e CaP MultiTask 300epi 10fps This is a public LeRobot v3.0 multi-task dataset for UR7e Code-as-Policies manipulation. It merges three 100-episode 10fps datasets into one 300-episode training corpus while preserving the original numeric observations, actions, and task labels. The final uploaded copy stores all videos as 10fps H.264 MP4 files. Dataset Summary Robot: ur7e Format: LeRobot v3.0 FPS: 10 Episodes: 300 Frames: 152,926 Cameras:… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/UR7e_CaP_MultiTask_300epi_10fps.tabularrobotics100K<n<1M0 likes108 downloads5mo agoHugging Face15chronbmm /sanskrit-multitasktext1M<n<10M0 likes80 downloads2y agoHugging Face16bigstupidhats /dynasample_multitasks_cleantabular1M<n<10M0 likes76 downloads2y agoHugging Face17chronbmm /sanskrit-multitask-devanagaritext1M<n<10M0 likes71 downloads2y agoHugging Face18Cache-SCA /SO101-cap_multitask_7tasks_700epi_10fpsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 700, "total_frames": 362160, "total_tasks": 7, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:700" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/SO101-cap_multitask_7tasks_700epi_10fps.tabularrobotics100K<n<1M0 likes71 downloads5mo agoHugging Face19TAUR-dev /11_9_25_STAR_multitask_sftdata_cd34_lm3_lc4_a4text10K<n<100K0 likes69 downloads11mo agoHugging Face20narendarcodes /Telugu-MultiTask-Instruct-77K Telugu MultiTask Instruct 77K — Adaption AutoScientist Challenge Dataset Powered by Adaptive Data — Adaption Labs Dataset Description A large-scale, multi-task Telugu instruction-tuning dataset combining 77,653 rows from 7 open-source Telugu NLP collections. Covers diverse tasks including news summarization, QA, creative writing, translation, and general instruction following — all processed through the Adaption Labs AutoScientist platform for quality… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/Telugu-MultiTask-Instruct-77K.textquestion-answering10K<n<100K1 likes68 downloads3mo agoHugging Face21rahuldshetty /satellite-multitask-omni 🛰️ Satellite Multi-Task Omni Dataset A unified, multi-task satellite/aerial imaging dataset designed for training omni-models that work with image+text as both input and output modalities. All data is converted to a consistent ChatML conversational format. 📊 Dataset Overview Metric Value Total Samples 34,894 Train / Val / Test 31,404 / 1,744 / 1,746 Tasks 9 distinct task types Sources 10 source datasets Format ChatML conversations + images… See the full description on the dataset page: https://huggingface.co/datasets/rahuldshetty/satellite-multitask-omni.imageimage-classification10K<n<100K1 likes64 downloads6mo agoHugging Face22kaustubhg73 /multilingual-multitask-refusal Multilingual Multitask Refusal A multilingual, multitask prompt dataset for analysing refusal behaviour across language, task wrapper, and harmful / harmless labels. English seed spans and labels come from the previous Multitask Multilingual Refusal dataset. Non-English content is produced with Google Sheets GOOGLETRANSLATE. Task instructions are language-localized via templates_localized.json. Rows 211,320 English seeds 1,761 Languages 15 Tasks 8 Product 1… See the full description on the dataset page: https://huggingface.co/datasets/kaustubhg73/multilingual-multitask-refusal.texttext-generation100K<n<1M0 likes62 downloads1mo agoHugging Face23Cache-SCA /UR7e_CaP_MultiTask_300epi_10fps_state_tplus1_action UR7e CaP MultiTask 300epi 10fps This is a public LeRobot v3.0 multi-task dataset for UR7e Code-as-Policies manipulation. It merges three 100-episode 10fps datasets into one 300-episode training corpus while preserving the original numeric observations, actions, and task labels. The final uploaded copy stores all videos as 10fps H.264 MP4 files. Dataset Summary Robot: ur7e Format: LeRobot v3.0 FPS: 10 Episodes: 300 Frames: 152,926 Cameras:… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/UR7e_CaP_MultiTask_300epi_10fps_state_tplus1_action.tabularrobotics100K<n<1M0 likes61 downloads5mo agoHugging Face24wcarvalho /MultitaskPreplay_jaxmaze_human_dftabular10K<n<100K0 likes60 downloads1y agoHugging Face25Nini0la /edgeimci-beta0-1k-multitask-enriched-2258-v1 EdgeIMCI Beta0-1K Multitask Enriched 2258 Dataset summary This is the public, message-normalized publication of the dataset used to fine-tune the selected EdgeIMCI 4B-alpha-second research checkpoint from the pinned Qwen/Qwen3-4B base model. It combines structured extraction, routing, bounded clinical-response, project self-knowledge, scope-safety, and proposition/negation examples for the EdgeIMCI sick-child assessment workflow. The dataset is research evidence… See the full description on the dataset page: https://huggingface.co/datasets/Nini0la/edgeimci-beta0-1k-multitask-enriched-2258-v1.texttext-generation1K<n<10K0 likes59 downloads18d agoHugging Face26TaskPuppyAI /lunamax-multitask-programming-1000 LunaMax Multitask Programming 1000 A 1,000-record synthetic multitask programming dataset generated with ChatGPT LunaMax. The recovered dataset combines code review, implementation, bug and severity classification, and strict output-contract tasks across multiple programming languages. The historical source shards were reviewed with ChatGPT 5.6 Sol High according to dataset creator confirmation. During Hugging Face publication preparation, all 1,000 records received a new… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-multitask-programming-1000.text1K<n<10K0 likes53 downloads1mo agoHugging Face27haoxianc /viveksingh2400_multitask-nlp OpenMultiTask NLP Dataset Mirror of the Kaggle dataset viveksingh2400/multitask-nlp by Vivek Singh, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data. A synthetic, logically constrained dataset for multi-task NLP including sentimen License MIT License, Copyright (c) Vivek Singh. The full license text is in LICENSE and applies to all files in this repository. Original description… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/viveksingh2400_multitask-nlp.text10K<n<100K0 likes52 downloads4d agoHugging Face28riddickz /multitask_v3_coding_11ktabular10K<n<100K0 likes48 downloads1y agoHugging Face29riddickz /multitask_v3_math_11ktabular10K<n<100K0 likes47 downloads1y agoHugging Face30Kowsher /multitask_textqa_benchmarkThis dataset is a part of . 🍈 MMT-47: Multimodal Multi-Task Benchmark 47 Tasks · 7 Categories · 3 Modalities (Image, Video, Text) Cite our ICML-2026 paper for this dataset @article{kowsher2026lime, title={LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning}, author={Kowsher, Md and Mansoor, Haris and Prottasha, Nusrat Jahan and Garibay, Ozlem and Zhu, Victor and Ji, Zhengping and Chen, Chen}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/Kowsher/multitask_textqa_benchmark.text10K<n<100K0 likes47 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.