Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /LAION-Audio-300Maudio100M<n<1B74 likes8.9k downloads2y agoHugging Face02mitermix /audiosnippetsaudio1M<n<10M6 likes4k downloads2y agoHugging Face03laion /laion-audio-previewaudio1M<n<10M11 likes2.6k downloads2y agoHugging Face04mitermix /audiosnippets_small_with_detailed_annotationaudio100K<n<1M1 likes1.1k downloads2y agoHugging Face05seungheondoh /cmd-audio-dumpaudio10K<n<100K1 likes998 downloads1y agoHugging Face06mitermix /audiosnippets_long_2_5Maudio1M<n<10M3 likes965 downloads2y agoHugging Face07mitermix /audiosnippets_small_with_detailed_annotation2audio1M<n<10M1 likes849 downloads2y agoHugging Face08laion /timbre-audio-caption-pairsaudio100K<n<1M2 likes783 downloads10mo agoHugging Face09mkrausio /audiosnippets-cleaned Dataset Summary This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications. Processing Details Transcriptions and broken characters were removed. All MP3 audio files were resampled to 16kHz for consistency. The accompanying JSON metadata was made consistent. Entries with… See the full description on the dataset page: https://huggingface.co/datasets/mkrausio/audiosnippets-cleaned.audio1M<n<10M3 likes684 downloads2y agoHugging Face10shenyunhang /AudioQA-1Maudio1M<n<10M2 likes674 downloads1y agoHugging Face11sleeping-ai /MemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata. We share these files as-part of research initiative. audio100K<n<1M0 likes630 downloads1y agoHugging Face12ssz1111 /SpokenWOZ-Train-Audioaudio1K<n<10K1 likes427 downloads9mo agoHugging Face13mitermix /audiosnippets_long_1Maudio100K<n<1M0 likes401 downloads2y agoHugging Face14Zhaowc /AudioCapsaudio100K<n<1M2 likes349 downloads1y agoHugging Face15acul3 /Audiobook_Noice_V2audio10K<n<100K0 likes292 downloads2y agoHugging Face16mitermix /audioset-with-grounded-captionsaudio1M<n<10M4 likes256 downloads1y agoHugging Face17laion /talent_plus_rl_groups_of_50_with_audiobox_scoresaudio1M<n<10M0 likes244 downloads10mo agoHugging Face18freococo /rohingya_asr_audioThis is the first public Rohingya language ASR dataset in AI history. Overview This dataset contains broadcast audio recordings from the Voice of America (VOA) Rohingya Service. Each file represents a daily news segment, typically 30 minutes in length, automatically segmented into chunks of 5–15 seconds for use in self-supervised ASR, pretraining, language identification, and more. The content was aired publicly as part of VOA’s Rohingya-language radio program and is therefore… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rohingya_asr_audio.audioautomatic-speech-recognition100K<n<1M2 likes204 downloads1y agoHugging Face19sheng22213 /speech_text-tts_audioaudio10K<n<100K0 likes193 downloads1y agoHugging Face20freococo /voa_myanmar_asr_audio_1 📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology. Overview This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.audioautomatic-speech-recognition1M<n<10M1 likes165 downloads1y agoHugging Face21overfitprolabse /subjective_audio_quality Balanced Perceptual Audio Quality Dataset Dataset Summary This is a large-scale, balanced dataset designed for training models for perceptual audio quality assessment. It consists of 612,020 examples, each containing a pair of 1-second audio clips: a high-quality original and a degraded version processed by various audio codecs. Each pair is accompanied by a perceptual quality score (ranging from 0.0 to 1.0) generated by visqol-like algorithms. The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/overfitprolabse/subjective_audio_quality.audio100K<n<1M1 likes141 downloads11mo agoHugging Face22TTS-AGI /balanced-audio-snippets-40x3k-DACVAEtext100K<n<1M0 likes138 downloads7mo agoHugging Face23Darknsu /mead_hdtf_400_merge_video_audio_frames_onlyimage1M<n<10M0 likes132 downloads4mo agoHugging Face24cmeraki /audiofolder_webdatasetaudio100K<n<1M0 likes98 downloads2y agoHugging Face25TTS-AGI /enhanced-audiosnippets-DACVAEtext1M<n<10M1 likes89 downloads7mo agoHugging Face26Helios1208 /Kling-Audio-Eval-cachetext10K<n<100K1 likes72 downloads5mo agoHugging Face27laion /audioset-with-captionsaudio1M<n<10M2 likes69 downloads11mo agoHugging Face28krishnakalyan3 /laion-audio-preview-splitaudio1M<n<10M2 likes64 downloads2y agoHugging Face29ReopenAI /Zhihu-KOL-Aug-Audio本数据集基于https://huggingface.co/datasets/wangrui6/Zhihu-KOL作为种子问题,使用Qwen2.5-72B-Instruct-GPTQ-Int4继续生成更多轮次的问题。 然后使用Qwen2.5-72B-Instruct-GPTQ-Int4生成问题的答案(每轮答案生成都会将之前的问题和答案当作上下文,确保当前的答案和历史相关)。见sharegpt.json文件。 问题使用cosyvoice生成对应音频。audio_part0-5.tar.gz是问题音频的压缩包。 audio100K<n<1M1 likes52 downloads1y agoHugging Face30niloy629 /imtalker-helium-audio-8s-backup IMTalker Helium + Audio 8s Backup Tar archive backup made before closing pod on 2026-05-13. Files: audio_hdtf_tf_helium_25fps.tar: 8s Helium-derived 25fps audio features audio_hdtf_tf_helium_25fps_meta.tar: metadata for Helium features audio_hdtf_tf_adapter768.tar: 8s adapter/wav2vec-style 768 audio features audio_hdtf_tf_adapter768_meta.tar: metadata for adapter768 features checkpoints/fp32_1layer_8s_best.pt checkpoints/fp32_12layer_8s_best.pt… See the full description on the dataset page: https://huggingface.co/datasets/niloy629/imtalker-helium-audio-8s-backup.text10K<n<100K0 likes52 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.