Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_3 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3" More Information needed audio10K<n<100K0 likes390 downloads3y agoHugging Face02Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_2 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2" More Information needed audio10K<n<100K0 likes310 downloads3y agoHugging Face03Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_6 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6" More Information needed audio10K<n<100K0 likes136 downloads3y agoHugging Face04Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_4 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4" More Information needed audio10K<n<100K0 likes128 downloads3y agoHugging Face05Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_1 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1" More Information needed audio10K<n<100K0 likes119 downloads3y agoHugging Face06Jayem-11 /mozilla_commonvoice_hackathon_preprocessed_train_batch_5 Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5" More Information needed audio10K<n<100K0 likes110 downloads3y agoHugging Face07build-small-hackathon /lolaby-traces Lolaby — generation traces Pipeline traces from Lolaby, an AI-powered lullaby generator built for the Build Small Hackathon 2026 (Backyard AI track). Each trace is a complete witness of one end-to-end generation: every input the user gave, every model that ran, every prompt and raw output, every timing measurement, and the final audio. Published under CC0 so anyone can study, replay, or remix the pipeline. What's in a trace Each subfolder is one generation. Files:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lolaby-traces.audiotext-to-audion<1K0 likes73 downloads4mo agoHugging Face08yasalma /tat_hackathon_asr Hackathon Tatar ASR Dataset Summary Hackathon Tatar ASR is a speech dataset distributed during the "Татар.Бу Хакатон" (Tatar.Bu Hackathon) held in Tatarstan in May 2024. This dataset likely consists of newly collected crowdsourced recordings created after the last release of TatSC (Tatar Speech Corpus), although some intersections with TatSC might be present. While TatSC contains 269.1 hours of transcribed speech with 271,914 utterances, this hackathon dataset comprises… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tat_hackathon_asr.audioautomatic-speech-recognition10K<n<100K0 likes28 downloads1y agoHugging Face09DORI-SRKW /orcasound_hackathon2024audioaudio-classification1K<n<10K0 likes22 downloads2y agoHugging Face10evie-8 /kinyarwanda-hackathongatedaudio100K<n<1M0 likes10 downloads1y agoHugging Face11evie-8 /kinyarwanda-speech-hackathongated 📚 Kinyarwanda ASR Dataset This dataset contains transcribed Kinyarwanda audio, designed to support training and evaluation of Automatic Speech Recognition (ASR) systems. It is part of a study on how varying training data volumes affect model performance using Whisper-large-v3. 📂 Data Overview The full dataset consists of approximately 263,000 audio samples covering 5 key domains: 🏥 Health 🏛️ Government 💰 Financial Services 🎓 Education 🌾 Agriculture To… See the full description on the dataset page: https://huggingface.co/datasets/evie-8/kinyarwanda-speech-hackathon.audio100K<n<1M0 likes10 downloads1y agoHugging Face12jq /kinyarwanda-speech-hackathongatedaudio100K<n<1M0 likes6 downloads1y agoHugging Face13ofarrelle /higgs-hackathon-2025audion<1K0 likes6 downloads1y agoHugging Face14build-small-hackathon /midnight-static-assetsaudion<1K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.