Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface-course /audio-course-imagesimagen<1K0 likes10k downloads3y agoHugging Face02GrunCrow /BIRDeep_AudioAnnotations BIRDeep Audio Annotations The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification. The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/GrunCrow/BIRDeep_AudioAnnotations.audioaudio-classificationn<1K2 likes1.3k downloads11mo agoHugging Face03tardellirs /enem-audiodescricao Audiodescrição profissional de imagens do ENEM: corpus em português para acessibilidade e avaliação de modelos de visão Professional audio descriptions of ENEM exam figures: a Brazilian Portuguese corpus for accessibility research and vision-language model evaluation. Corpus de audiodescrições escritas por profissionais para as figuras do ENEM, extraídas dos cadernos "ledor" que o INEP publica para participantes com deficiência visual, alinhadas à figura, ao enunciado, às… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/enem-audiodescricao.imageimage-to-textn<1K0 likes1.2k downloads1mo agoHugging Face04raymondt /tien_lo_audio426imagen<1K0 likes1k downloads1d agoHugging Face05jiaheillu /sovits_audio_preview 预览. 简体中文| English| 日本語 本仓库用于预览so-vits-svc-4.0训练出的各种语音模型的效果,点击角色名自动跳转对应训练参数。 推荐用谷歌浏览器,其他浏览器可能无法正确加载预览的音频。 正常说话的音色转换较为准确,歌曲包含较广的音域且bgm和声等难以去除干净,效果有所折扣。 有推荐的歌想要转换听听效果,或者其他内容建议,点我发起讨论 下面是预览音频,上下左右滑动可以看到全部 角色名 角色原声A 被转换人声BA音色替换B A音色翻唱(点击直接下载) 散兵 夢で会えたら 胡桃 ......... ......... moonlight shadow, 云烟成雨… See the full description on the dataset page: https://huggingface.co/datasets/jiaheillu/sovits_audio_preview.audion<1K8 likes926 downloads3y agoHugging Face06yamahaymh /karaoke-tmp-audio-pubaudion<1K0 likes857 downloads13h agoHugging Face07raymondt /van_hai_audioimagen<1K0 likes531 downloads17d agoHugging Face08raymondt /cuu_tieu_audioimagen<1K0 likes500 downloads18d agoHugging Face09TenzinL /BIRDeep_AudioAnnotations BIRDeep Audio Annotations The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification. The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/TenzinL/BIRDeep_AudioAnnotations.audioaudio-classificationn<1K0 likes496 downloads8mo agoHugging Face10lmms-lab-audio /Omni_Bench_fixaudio1K<n<10K1 likes491 downloads2y agoHugging Face11raymondt /dao_tam_audioimagen<1K0 likes489 downloads17d agoHugging Face12hammoualiyoucef20 /quran-audioaudion<1K0 likes474 downloads11d agoHugging Face13raymondt /huyen_khong_audioimagen<1K0 likes472 downloads18d agoHugging Face14raymondt /huyen_vu_audioimagen<1K0 likes413 downloads17d agoHugging Face15hf-audio /gradient_accumulation_exampleimagen<1K0 likes393 downloads2y agoHugging Face16RyanWW /audiobench_rendertextaudio10K<n<100K0 likes351 downloads1y agoHugging Face17trentmkelly /wwii_audio_transcribed WWII Audio with Transcripts 993 World War II-era recordings from the Internet Archive WWII audio collection, with machine-generated transcripts: 1944: 558 recordings 1945: 435 recordings Total: approximately 220 hours of audio; 6.78 GB including alternate audio formats, transcripts, and archive images. The dataset viewer pairs each recording with playable audio and its full transcript. Transcripts were generated with Microsoft MAI Transcribe 2 and may contain errors or be… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/wwii_audio_transcribed.audioautomatic-speech-recognitionn<1K0 likes273 downloads18d agoHugging Face18teticio /audio-diffusion-1024Over 20,000 256x256 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 1024 y_res = 1024 sample_rate = 44100 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K0 likes236 downloads4y agoHugging Face19raymondt /kiem_y_audioimagen<1K0 likes180 downloads17d agoHugging Face20raymondt /huyen_co_audioimagen<1K0 likes177 downloads17d agoHugging Face21raymondt /kiem_vu_audioimagen<1K0 likes174 downloads17d agoHugging Face22ziggylott /audiobooksaudion<1K0 likes150 downloads3mo agoHugging Face23AhunInteligence /Amharic_Audio_and_Spectrograms Amharic Audio Spectrogram Dataset Dataset Info Total samples in full dataset: 662,611 Samples in this preview: 1,000 Audio duration: 2.49 ± 1.60 seconds Sample rate: 16kHz Spectrogram dimensions: 80 mel bins × variable time steps Sample Data Audio Sample Spectrogram License Apache 2.0 audio1K<n<10K0 likes136 downloads1y agoHugging Face24Quazitron420 /video-dataset-audio_dataset Video Dataset - audio_dataset Dataset Description This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips. Combined from tasks: task06, task07, task08 Dataset Structure frames/ — extracted frames (first frame from each segment) segments/ — video clips for each annotation interval annotations/ — original JSON annotation transcriptions/ — transcription files… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-audio_dataset.imageimage-classificationn<1K0 likes132 downloads11mo agoHugging Face25Darknsu /mead_hdtf_400_merge_video_audio_frames_onlyimage1M<n<10M0 likes132 downloads4mo agoHugging Face26raymondt /kiem_tam_audioimagen<1K0 likes118 downloads17d agoHugging Face27teticio /audio-diffusion-512Over 20,000 512x512 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 512 y_res = 512 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K2 likes117 downloads3y agoHugging Face282bcountingu /NTU60-AUDIOimagen<1K0 likes115 downloads8mo agoHugging Face29Rapidata /kiki-bouba-audio-20k 🔊 Kiki–Bouba, Spoken Aloud (20k Global Responses) Dataset Summary This dataset is the audio companion to Rapidata/psychology-association-kiki-bouba-etc. In the original dataset, respondents read the question "Which one is called 'Kiki'?" as written text. Here, respondents instead hear the word spoken aloud — the task shows the same two shapes (a rounded blob and a spiky star) while a short audio clip of "kiki" or "bouba" plays as context. The annotator UI… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/kiki-bouba-audio-20k.audion<1K2 likes111 downloads2mo agoHugging Face30danjacobellis /audiosetimage10K<n<100K0 likes110 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.