Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ExylosAi /egocentric-vr-capture-20h-multimodal-sample Egocentric VR Capture — 20-Hour Multimodal Inspection Sample 195 real-world task episodes / 2,283,482 frames / 21.14 delivered hours captured with consumer VR hardware. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package. This publicly accessible 20-hour-scale dataset is produced by the EXYLOS real-world data pipeline. Files and the Dataset Viewer can be accessed without individual approval;… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/egocentric-vr-capture-20h-multimodal-sample.tabularrobotics1M<n<10M0 likes1.6k downloads18d agoHugging Face02ShulongZhang /Multimodal_Fish_Feeding_Intensityaudio1K<n<10K0 likes1.2k downloads1y agoHugging Face03Hariprasath5128 /marine-animals-multimodal-dataset Marine Animals Multimodal Dataset 🐋 A comprehensive multimodal dataset combining audio recordings and images of 32 marine species. Dataset Summary Total samples: 24,911 Species: 32 Audio files: 1,357 unique recordings Images: 581 (309 matched + 272 from iNaturalist) Features species (string): Species name label (int32): Numeric label (0–31) audio (Audio): Audio recording of the species image (Image): Species image image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.audioaudio-classification10K<n<100K0 likes846 downloads10mo agoHugging Face04Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes751 downloads4mo agoHugging Face05mdimamhosen /pd-voice-full-multimodal-dataset Parkinson Voice — Full Multimodal Dataset Complete Parkinson’s vs healthy voice package for classification and explainable Mel reasoning research (EDGE). Not Mel-only: raw audio, 10 visual modalities, feature CSVs, plus Gemma reasoning traces for Mel. Contents Path Description audio/ Waveform clips (healthy / parkinsons), 1134 files images/mel/ Mel spectrograms images/spectrogram/ Linear spectrograms images/mfcc/ MFCC maps images/delta_mfcc/… See the full description on the dataset page: https://huggingface.co/datasets/mdimamhosen/pd-voice-full-multimodal-dataset.imageaudio-classification10K<n<100K0 likes532 downloads1mo agoHugging Face06SilverAvocado /Silver-Multimodal-Dataset Dataset Overview The dataset is designed to support the development of machine learning models for detecting daily activities, violence, and fall down scenarios from combined audio and video sources. The preprocessing pipeline leverages audio feature extraction, human keypoint detection, and relative positional encoding to generate a unified representation for training and inference. Classes: 0: Daily - Normal indoor activities 1: Violence - Aggressive behaviors 2: Fall Down -… See the full description on the dataset page: https://huggingface.co/datasets/SilverAvocado/Silver-Multimodal-Dataset.videoobject-detection100B<n<1T0 likes338 downloads2y agoHugging Face07IceKhoffi /chicken-health-behavior-multimodal Chicken Health and Behavior Multimodal Dataset (chicken-health-behavior-multimodal) This dataset provides a comprehensive collection of visual and audio data from chicken farms, specifically designed for early detection of chicken health issues and anomalous behaviors. It aims to support the development of intelligent monitoring systems that can mitigate significant economic losses caused by poultry disease, with a particular focus on future applications. Why this… See the full description on the dataset page: https://huggingface.co/datasets/IceKhoffi/chicken-health-behavior-multimodal.audioobject-detectionn<1K4 likes309 downloads1y agoHugging Face08PRAIG /grandstaff-grandstaff-multimodalaudio10K<n<100K1 likes233 downloads1y agoHugging Face09cjerzak /MultimodalMathBenchmarks MultimodalMathBenchmarks This repository contains the datasets for the paper Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (ACL Findings 2026). It covers the public benchmark datasets and their modality assets (text, images, and audio) used to evaluate the arithmetic capabilities of multimodal LLMs. Canonical Upload Manifest HF path Local source Count Purpose SharedMultimodalGrid.csv SavedData/SharedMultimodalGrid.csv… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/MultimodalMathBenchmarks.audioimage-text-to-text10K<n<100K0 likes218 downloads2mo agoHugging Face10IntelLabs /Intel_Robotic_Welding_Multimodal_Datasetgated Dataset Card for the Intel Robotic Welding Multimodal Dataset This dataset was collected to enable multimodal welding defect detection research. The dataset contains over 4000 annotated samples and was collected in an automotive production floor setting in collaboration with a supplier with access to such facilities. Each sample contains a video, associated audio, a time-series from welding sensors, and five post-weld images for a particular weld. A separately licensed… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/Intel_Robotic_Welding_Multimodal_Dataset.audio67 likes173 downloads1y agoHugging Face11J017athan /Multimodal-Yue-Benchmark Multimodal Yue Benchmark Cantonese audio + text benchmark derived from BillBao/Yue-Benchmark (Yue-GSM8K & Yue-MMLU-style tasks). We kept the same task content in Cantonese and added TTS for three Cantonese speakers (hiugaai, hiumaan, wanlung). Subsets & splits Config name Task Speaker Splits mmlu_* multiple-choice (MMLU-style) per speaker train, test gsm8k_* math word problems (GSM8K-style) per speaker train, test Example: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/J017athan/Multimodal-Yue-Benchmark.audio10K<n<100K0 likes159 downloads7mo agoHugging Face12introvoyz042 /egocentric-vr-capture-1h-multimodal-sample Egocentric VR Capture — 1-Hour Multimodal Inspection Sample 13 real-world task episodes / 108,029 frames / approximately 60 minutes captured with Meta Quest 3. Each episode combines egocentric RGB and audio with synchronized headset, camera, body, and hand tracking in a LeRobot v3-style package. This publicly accessible dataset is an inspection slice produced by the EXYLOS real-world data pipeline. It demonstrates capture quality, synchronization, schema, and QA metadata… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz042/egocentric-vr-capture-1h-multimodal-sample.tabularrobotics100K<n<1M0 likes148 downloads1mo agoHugging Face13fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes148 downloads22d agoHugging Face14fluid-concepts /multimodal-expert-instruction-samplesgated Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside. ▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.audion<1K1 likes134 downloads22d agoHugging Face15Danetuk /chicken-health-behavior-multimodal Chicken Health and Behavior Multimodal Dataset (chicken-health-behavior-multimodal) This dataset provides a comprehensive collection of visual and audio data from chicken farms, specifically designed for early detection of chicken health issues and anomalous behaviors. It aims to support the development of intelligent monitoring systems that can mitigate significant economic losses caused by poultry disease, with a particular focus on future applications. Why this… See the full description on the dataset page: https://huggingface.co/datasets/Danetuk/chicken-health-behavior-multimodal.audioobject-detectionn<1K0 likes130 downloads3mo agoHugging Face16audibeal74 /panta_instruct_multi_modal_v1 Panta Instruct Multi-Modal v1 Dataset d'instructions multimodal en français : chaque exemple associe une question (texte + parole + pictogrammes) à une réponse (texte + pictogrammes). Colonnes Colonne Type Description audio Audio (24 kHz, mono) Enregistrement de la question (text_input) text_input string Question / instruction text_output string Réponse pictos_input list[string] Identifiants des pictogrammes de la question pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.audioautomatic-speech-recognition10K<n<100K0 likes115 downloads17d agoHugging Face17geoffbremneraudio /Geoff_Bremner_Multimodal_Music_Corpus_SAMPLE Geoff Bremner Multimodal Music Corpus — Sample Release This is a single-track sample from the Geoff Bremner Multimodal Music Corpus, a growing, research-grade, commercially licensable dataset of 100% original music — written, recorded, and produced entirely by one artist . If this sample meets your needs - please contact me directly for more Geoff Bremner https://linktr.ee/gbaudio License This dataset is released under CC BY-NC 4.0… See the full description on the dataset page: https://huggingface.co/datasets/geoffbremneraudio/Geoff_Bremner_Multimodal_Music_Corpus_SAMPLE.audion<1K2 likes111 downloads3mo agoHugging Face18FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes104 downloads5mo agoHugging Face19mrzjy /multimodal-genshin-impact Genshin Impact Fandom Wiki Multimodal Dataset Github repo here Description This dataset is a comprehensive collection of 22,162 fandom wiki pages for the popular game Genshin Impact. The dataset includes markdown-formatted English content from the wiki, featuring interleaved text, as well as image, video, and audio file links. Additionally, the associated multimodal files (images, videos, and audio) have been downloaded and organized to facilitate the multimodal dataset… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/multimodal-genshin-impact.audio10K<n<100K4 likes96 downloads2y agoHugging Face20PRAIG /chopin-grandstaff-multimodalaudio1K<n<10K0 likes85 downloads1y agoHugging Face21opendatalab /WanJuanSiLu-Multimodal-5Languages WanJuan·SiLu Multimodal Multilingual Corpus 🌏Dataset Introduction The newly upgraded "Wanjuan·Silk Road Multimodal Corpus" brings the following three core improvements: The number of languages has been significantly expanded: Based on the five open-source languages ​​of "Wanjuan·Silk Road", namely Arabic, Russian, Korean, Vietnamese, and Thai, "Wanjuan·Silk Road Multimodal" has added three scarce corpus data of Serbian, Hungarian, and Czech, and uses the above… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuanSiLu-Multimodal-5Languages.audio100K<n<1M4 likes82 downloads1y agoHugging Face22nativemind /kene_multimodal_gift Kene Multimodal Gift Dataset (Enhanced with Ethnic Languages) Описание Мультимодальный духовный датасет с ИКАРОС на испанском, Джив Джаго на хинди и языками народностей России, СНГ и Украины. Обновления ✅ Добавлены ИКАРОС на испанском языке ✅ Добавлен Джив Джаго на хинди ✅ НОВОЕ: Добавлены языки народностей России, СНГ и Украины ✅ Улучшены мультимодальные данные ✅ Расширена поддержка языков до 50 примеров Языки Духовные языки Русский:… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/kene_multimodal_gift.audion<1K1 likes68 downloads1y agoHugging Face23Bisher /ClArTTS-multimodalaudio1K<n<10K0 likes66 downloads1y agoHugging Face24danielrosehill /multimodal-ai-taxonomy Multimodal AI Taxonomy A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities. Dataset Description This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to: Navigate the complex landscape of multimodal AI models Filter models by specific input/output modality combinations Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.textothern<1K0 likes65 downloads1y agoHugging Face25ffurfaro /keep-it-simple-multimodal keep-it-simple-multimodal A mini, standalone multimodal dataset: image+caption, audio+caption, video+caption, lidar, IMU, and optimal-control state/action pairs. Companion to keep-it-simple (text), built to feed KairosPretrainingDataset in kairos. Structure One generic schema for every row — no per-modality columns, no fixed shape/dtype assumptions: Column Type Description modality string image_caption | audio_caption | video_caption | lidar | imu |… See the full description on the dataset page: https://huggingface.co/datasets/ffurfaro/keep-it-simple-multimodal.textother1K<n<10K0 likes60 downloads2mo agoHugging Face26nativemind /mozgach_multimodal_extraaudion<1K0 likes59 downloads1y agoHugging Face27woojun-jung /multimodal-guide-using-huggingfaceaudion<1K0 likes54 downloads8mo agoHugging Face28lv12 /MultiModalDataset Dataset Card for MultiModal Dataset Dataset Description Dataset Summary MultiModal Dataset is a curated collection of 85,000 samples spanning three modalities: text, images, and audio. It combines high-quality web content, image-caption pairs from COCO 2017, and audio samples from AudioSet to enable comprehensive multimodal model training and evaluation. The dataset is organized into three subsets: fineweb: 37,500 high-quality web text samples (>8… See the full description on the dataset page: https://huggingface.co/datasets/lv12/MultiModalDataset.imagetext-generation10K<n<100K0 likes52 downloads2mo agoHugging Face29monish-73 /marine-animals-multimodalaudio10K<n<100K0 likes45 downloads10mo agoHugging Face30Maisum-Abbas-123 /UMED-Urdu-Multimodal-Emotion-Datasetaudio1K<n<10K1 likes44 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.