datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION-Audio-300Maudiosnippetslaion-audio-previewaudiosnippets_small_with_detailed_annotationcmd-audio-dumpaudiosnippets_long_2_5Maudiosnippets_small_with_detailed_annotation2timbre-audio-caption-pairsaudiosnippets-cleaned
Dataset Summary
This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications.
Processing Details
Transcriptions and broken characters were removed.
All MP3 audio files were resampled to 16kHz for consistency.
The accompanying JSON metadata was made consistent.
Entries with… See the full description on the dataset page: https://huggingface.co/datasets/mkrausio/audiosnippets-cleaned.AudioQA-1MMemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata.
We share these files as-part of research initiative.
SpokenWOZ-Train-Audioaudiosnippets_long_1MAudioCapsAudiobook_Noice_V2audioset-with-grounded-captionstalent_plus_rl_groups_of_50_with_audiobox_scoresrohingya_asr_audioThis is the first public Rohingya language ASR dataset in AI history.
Overview
This dataset contains broadcast audio recordings from the Voice of America (VOA) Rohingya Service. Each file represents a daily news segment, typically 30 minutes in length, automatically segmented into chunks of 5–15 seconds for use in self-supervised ASR, pretraining, language identification, and more.
The content was aired publicly as part of VOA’s Rohingya-language radio program and is therefore… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rohingya_asr_audio.speech_text-tts_audiovoa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.subjective_audio_quality
Balanced Perceptual Audio Quality Dataset
Dataset Summary
This is a large-scale, balanced dataset designed for training models for perceptual audio quality assessment. It consists of 612,020 examples, each containing a pair of 1-second audio clips: a high-quality original and a degraded version processed by various audio codecs. Each pair is accompanied by a perceptual quality score (ranging from 0.0 to 1.0) generated by visqol-like algorithms.
The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/overfitprolabse/subjective_audio_quality.balanced-audio-snippets-40x3k-DACVAEmead_hdtf_400_merge_video_audio_frames_onlyaudiofolder_webdatasetenhanced-audiosnippets-DACVAEKling-Audio-Eval-cacheaudioset-with-captionslaion-audio-preview-splitZhihu-KOL-Aug-Audio本数据集基于https://huggingface.co/datasets/wangrui6/Zhihu-KOL作为种子问题,使用Qwen2.5-72B-Instruct-GPTQ-Int4继续生成更多轮次的问题。
然后使用Qwen2.5-72B-Instruct-GPTQ-Int4生成问题的答案(每轮答案生成都会将之前的问题和答案当作上下文,确保当前的答案和历史相关)。见sharegpt.json文件。
问题使用cosyvoice生成对应音频。audio_part0-5.tar.gz是问题音频的压缩包。
imtalker-helium-audio-8s-backup
IMTalker Helium + Audio 8s Backup
Tar archive backup made before closing pod on 2026-05-13.
Files:
audio_hdtf_tf_helium_25fps.tar: 8s Helium-derived 25fps audio features
audio_hdtf_tf_helium_25fps_meta.tar: metadata for Helium features
audio_hdtf_tf_adapter768.tar: 8s adapter/wav2vec-style 768 audio features
audio_hdtf_tf_adapter768_meta.tar: metadata for adapter768 features
checkpoints/fp32_1layer_8s_best.pt
checkpoints/fp32_12layer_8s_best.pt… See the full description on the dataset page: https://huggingface.co/datasets/niloy629/imtalker-helium-audio-8s-backup.
