Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MC7ever /hf-training-corpus HF Training Corpus Bulk-scraped multimodal training corpus: ~92,000 Hugging Face datasets streamed, normalized and stored as per-dataset Parquet files under data/. Every modality is captured: text, images (PNG bytes), audio (mono 16-bit WAV bytes), video (bytes, capped) and tabular/timeseries (serialized to text). Pipeline Live crawl from a 92,458-ID corpus list (see https://github.com/shadyuwugurl/hf-datasets-trees/blob/main/datasets.txt) streaming=True reads… See the full description on the dataset page: https://huggingface.co/datasets/MC7ever/hf-training-corpus.audiotext-generation1M<n<10M0 likes15k downloads4m agoHugging Face02ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes12k downloads8h agoHugging Face03facebook /2M-Flores-ASL 2M-Flores As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest sentences in the original flores200 dataset. To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded. The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time. The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.tabulartranslation1K<n<10K2 likes2.8k downloads2y agoHugging Face04FBK-MT /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.audioautomatic-speech-recognition1K<n<10K70 likes1.6k downloads3mo agoHugging Face05Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes1.2k downloads13d agoHugging Face06BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.1k downloads11mo agoHugging Face07Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes799 downloads4mo agoHugging Face08H-oliday /TeleEgo-Source TeleEgo-Source Source Videos and Time-Aligned Transcripts for TeleEgo Official source release forTeleEgo: Benchmarking Egocentric AI Assistants in the Wild Overview TeleEgo is a multimodal benchmark for evaluating egocentric AI assistants in realistic, long-duration settings. It contains recordings from five participants over three days and covers four broad themes: Work & Study, Lifestyle &… See the full description on the dataset page: https://huggingface.co/datasets/H-oliday/TeleEgo-Source.videovideo-text-to-textn<1K0 likes662 downloads1mo agoHugging Face09Shofo /shofo-talking-head-engated Shofo Talking Head Dataset (English) A curated dataset of ~10,000 talking-head videos (~186 hours, mean ~67s/clip), filtered for clean single-speaker framing and paired with time-aligned transcripts. Built and released by Shofo. This dataset is short-form social video, designed to support modern avatar, lip-sync, dubbing, and TTS work. Every clip is curated by a multi-stage pipeline (face/framing analysis, on-screen-text detection, object-occlusion detection, voice/face matching… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-talking-head-en.videovideo-classification1K<n<10K1 likes583 downloads3mo agoHugging Face10tugrulbayrak /Real-TurnTurk Real-TurnTurk English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.tabularaudio-classification100K<n<1M3 likes515 downloads15d agoHugging Face11humyn-labs /APAC-Egocentric-Residential-Voiceover APAC Egocentric Residential (with Voiceover) Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track. This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video. Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.tabularroboticsn<1K0 likes285 downloads2mo agoHugging Face12IndabaXSudan /Sudan-MM Sudan-MM: A Multimodal Dataset of Sudanese Arabic Sudan-MM is the first publicly available multimodal dataset for Sudanese Arabic (السودانية), a low-resource dialect with no prior paired image-caption, video-caption, or voice-caption data. It was produced through a competitive shared task held in 2025, where five teams collected and annotated media depicting everyday Sudanese life. Each item in the dataset pairs a visual or video recording with: a written caption in Modern Standard… See the full description on the dataset page: https://huggingface.co/datasets/IndabaXSudan/Sudan-MM.audioimage-to-text1K<n<10K2 likes241 downloads5mo agoHugging Face13Arozo /aiza-malagasy-administrative-asr AIZA — Malagasy/French Code-Switched Administrative Speech (Benchmark Set) 10 short, self-recorded, consented audio clips of Malagasy/French code-switched questions about Malagasy administrative procedures (lost ID card, birth certificate, passport, land title, etc.) — created as the benchmark audio for AIZA, a submission to the Sahara CodeSwitch Africa Main Challenge. Why this dataset exists Neither Sahara's published language list nor the official… See the full description on the dataset page: https://huggingface.co/datasets/Arozo/aiza-malagasy-administrative-asr.audioautomatic-speech-recognitionn<1K0 likes146 downloads20d agoHugging Face14ohsn /darija_yt_2026 darija_yt_2026 Partition upload generated automatically. Namespace: ohsn Repo: ohsn/darija_yt_2026 Video count: 3511 Duration hours: 1565.31 This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline. audioautomatic-speech-recognition1K<n<10K0 likes142 downloads1mo agoHugging Face15kadirnar /ai-researcher-roadmap-media AI Researcher Roadmap Media Optional video and subtitle assets for the AI Researcher Roadmap application. Repository layout manifest.json: file sizes and SHA-256 checksums used by the application. videos/<stem>.mp4: lecture video. subs/<stem>.<language>.vtt: subtitle tracks. subs/<stem>.asr.<language>.vtt: ASR-generated subtitle tracks. The application downloads only the selected lecture and its subtitle tracks. Files are cached locally and can be played offline… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/ai-researcher-roadmap-media.videoautomatic-speech-recognitionn<1K0 likes121 downloads3mo agoHugging Face16snorbyte /indic-audio-dialog-samplegated Dataset Card for Indic Dialog Sample Dataset Dataset Details Dataset Description The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.audioaudio-to-audio1K<n<10K1 likes78 downloads1y agoHugging Face17Fhrozen /CABankSakura CABank Japanese Sakura Corpus Susanne Miyata Department of Medical Sciences Aichi Shukotoku University smiyata@asu.aasa.ac.jp website: https://ca.talkbank.org/access/Sakura.html Important This data set is a copy from the original one located at https://ca.talkbank.org/access/Sakura.html. Details Participants: 31 Type of Study: xxx Location: Japan Media type: audio DOI: doi:10.21415/T5M90R Citation information Some citation here. In… See the full description on the dataset page: https://huggingface.co/datasets/Fhrozen/CABankSakura.videoaudio-classificationn<1K0 likes73 downloads4y agoHugging Face18Rendra86318 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes72 downloads10mo agoHugging Face19juliasdata /medical-audio-sample-brazilian-portuguese Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.audioautomatic-speech-recognitionn<1K1 likes70 downloads7mo agoHugging Face20tosinamuda /afromathvoices AfroMathVoices African-accented English speakers reading mathematical expressions aloud, paired with target LaTeX. For speech-to-LaTeX and mathematical speech recognition research. Hugging Face: tosinamuda/afromathvoices 20 recordings of 10 prompts, read by 2 speakers (Nigerian English, Kenyan English). Prompts and LaTeX come from the equations_test split of marsianin500/Speech2Latex. Audio is mono 48 kHz Opus in WebM, not transcoded. Version 0.1.0 is a pilot. It is too small… See the full description on the dataset page: https://huggingface.co/datasets/tosinamuda/afromathvoices.audioautomatic-speech-recognitionn<1K0 likes65 downloads7d agoHugging Face21kurianbenoy /Indic-subtitler-audio_evals Indic_audio_evals As part of this project. We are evaluating our performance of various ASR models as well in a benchmarking dataset, we have created in various languages. This benchmarking dataset is more alligned to real-world use-cases rather than having any academic datasets. About Dataset Dataset Link in HuggingFace: kurianbenoy/Indic-subtitler-audio_evals This dataset contains audio file in .wav format and video file in .mp4. The respective groundtruth will be… See the full description on the dataset page: https://huggingface.co/datasets/kurianbenoy/Indic-subtitler-audio_evals.audioautomatic-speech-recognitionn<1K2 likes56 downloads2y agoHugging Face22latentic /afromathvoices AfroMathVoices African-accented English speakers reading mathematical expressions aloud, paired with target LaTeX. For speech-to-LaTeX and mathematical speech recognition research. Hugging Face: latentic/afromathvoices 20 recordings of 10 prompts, read by 2 speakers (Nigerian English, Kenyan English). Prompts and LaTeX come from the equations_test split of marsianin500/Speech2Latex. Audio is mono 48 kHz Opus in WebM, not transcoded. Version 0.1.0 is a pilot. It is too small to… See the full description on the dataset page: https://huggingface.co/datasets/latentic/afromathvoices.audioautomatic-speech-recognitionn<1K0 likes47 downloads6d agoHugging Face23scubavoice /ibibio-efik-speech-corpus-sample Scuba Voice Dataset: Ibibio & Efik Sample (1 hour) Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria. This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.audioautomatic-speech-recognitionn<1K1 likes41 downloads3mo agoHugging Face24ZamAI-Pashto /zamai-pashto-video ZamAI Pashto Video ZamAI Pashto Video is a video understanding dataset scaffold for multilingual Afghan media research, with support for scene segmentation, subtitle alignment, and temporal event annotation. Dataset Summary The repository is structured for raw and segmented video assets, subtitle generation, action labels, and media metadata needed for temporal analysis workflows. Languages Pashto Dari English Modalities Video… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-video.textvideo-classificationn<1K0 likes38 downloads2mo agoHugging Face25octanove /moslagated Overview The MOSLA dataset ("MOSLA") is a longitudinal, multimodal, multilingual, and controlled dataset created by inviting participants to learn one of three target languages (Arabic, Spanish, and Chinese) from scratch over a span of two years, exclusively through online instruction, and recording every lesson using Zoom. The dataset is semi-automatically annotated with speaker/language IDs and transcripts by both human annotators and fine-tuned state-of-the-art speech models.… See the full description on the dataset page: https://huggingface.co/datasets/octanove/mosla.tabularautomatic-speech-recognition100K<n<1M5 likes35 downloads2y agoHugging Face26JacobLinCool /anime-2024gated Anime Video Dataset 2024 Overview This is an anime video dataset curated specifically for multimedia research purposes. It is sourced from anime series released in 2024 and includes both a complete combined file and four seasonal subsets (Winter, Spring, Summer, and Fall). Data Details Video Stream Codec: H.264 (High Profile) Format: YUV 4:2:0 (Progressive) Resolution: 640×360 (SAR 1:1, DAR 16:9) Frame Rate: 23.98 FPS Audio Stream Codec: AAC (LC) Sample… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/anime-2024.textvideo-classification1K<n<10K2 likes34 downloads11mo agoHugging Face27DigiGreen /AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya. The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts. The transcripts are: Generated by ASR models (for the purpose of benchmarking) Manual transcripts Time stamps Manual translations This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/AgricultureVideosTranscript.videotranslation1K<n<10K0 likes32 downloads2y agoHugging Face28alj68 /2M-Flores-ASL 2M-Flores As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest sentences in the original flores200 dataset. To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded. The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time. The… See the full description on the dataset page: https://huggingface.co/datasets/alj68/2M-Flores-ASL.tabulartranslation1K<n<10K0 likes32 downloads10mo agoHugging Face29vaishnavikedar4 /MCIF Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/vaishnavikedar4/MCIF.audioautomatic-speech-recognition1K<n<10K0 likes31 downloads9mo agoHugging Face30CGIAR /AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya. The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts. The transcripts are: Generated by ASR models (for the purpose of benchmarking) Manual transcripts Time stamps Manual translations This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosTranscript.videotranslation1K<n<10K0 likes29 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.