Team Ai
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.1k downloads8mo agoHugging Face02manassehzw /sna-dataset-annotated manassehzw/sna-dataset-annotated An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline. This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments. Why this annotated release exists The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.audioautomatic-speech-recognition10K<n<100K1 likes289 downloads2mo agoHugging Face03humair025 /Rasa-Annotated-25kHz Rasa-Annotated Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,102 Total Duration: 46.92 hours Average Duration: 6.47 seconds Duration Range: 0.31s - 45.34s Average Phonemes: 18.4 per sample Average Kanade Tokens: 530.7 per sample Global Embedding Dimension: 128 Gender Distribution Gender Count Female 12,583 Male 13,519 Style Distribution Style Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-25kHz.tabulartext-to-speech10K<n<100K0 likes220 downloads8mo agoHugging Face04manassehzw /sna-waxal-annotated-unlabeled Shona WAXAL annotated-unlabeled checkpoint This is a self-contained operational checkpoint for pseudo-labeling Shona ASR data. It contains 90,253 conservatively segmented FLAC clips (441.585 hours), but intentionally contains no transcripts. Fields transcription is intentionally empty. speaker_id is an approximate source-blind EOM cluster or unknown; speaker_clip_count is zero for unknown assignments. gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.audioautomatic-speech-recognition10K<n<100K0 likes215 downloads2mo agoHugging Face05Congo-digital-service /audios-lingala-annotatees-v2gated Annotated Lingala Audio — canonical corpus Annotated Lingala speech for open automatic speech recognition research and for fine-tuning speech models. This release is a full reconstruction of the corpus from its source recordings and annotations. It supersedes Congo-digital-service/audios-lingala-annotatees, which is deprecated — see Relationship to the previous release below. What this dataset contains Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.audioautomatic-speech-recognition10K<n<100K0 likes213 downloads29d agoHugging Face06Congo-digital-service /audios-lingala-annotateesgated Annotated Lingala Dataset – Full Version Description This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models. It includes: the original audio files (viewable directly in the Hugging Face viewer) text transcriptions Mel spectrograms tokenized labels Overall statistics Metric Value Total volume 5 h 0 min 18 s Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.audioautomatic-speech-recognition10K<n<100K0 likes195 downloads1mo agoHugging Face07projecte-aina /annotated_catalan_common_voice_v17This version of the Catalan sentences of the Common Voice corpus v17 includes metadata (gender and accent) for 263 speakers annotated by a team of experts.automatic-speech-recognition1M<n<10M1 likes107 downloads1y agoHugging Face08humair025 /Rasa-Annotated-V1 Rasa-Annotated Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,102 Total Duration: 46.92 hours Average Duration: 6.47 seconds Duration Range: 0.31s - 45.34s Average Phonemes: 18.4 per sample Average Kanade Tokens: 264.5 per sample Global Embedding Dimension: 128 Gender Distribution Gender Count Female 12,583 Male 13,519 Style Distribution Style Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-V1.tabulartext-to-speech10K<n<100K0 likes93 downloads8mo agoHugging Face09dikal /maleo-emotion-annotated Maleo Emotion Audio Dataset Indonesia (Annotated with Word-Level Timestamps) Dataset audio ucapan emosi bahasa Indonesia (Indonesian Speech Emotion Recognition) yang dilengkapi dengan transkripsi teks manual dan penanda waktu kata demi kata (word-level timestamps). Dataset ini ditujukan untuk penelitian dan pengembangan sistem Speech Emotion Recognition (SER), Automatic Speech Recognition (ASR), serta analisis multimodal ucapan emosional. Ringkasan Dataset… See the full description on the dataset page: https://huggingface.co/datasets/dikal/maleo-emotion-annotated.audioaudio-classificationn<1K0 likes70 downloads7d agoHugging Face10ebellob /annotated_catalan_common_voice_v17_cleaned_enhanced Processed Annotated Catalan Common Voice v17 (CleanUNet + FlashSR) Dataset Summary This dataset is a processed and enhanced version of: projecte-aina/annotated_catalan_common_voice_v17. Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove. However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/annotated_catalan_common_voice_v17_cleaned_enhanced.audiotext-to-speech100K<n<1M2 likes56 downloads6mo agoHugging Face11Congo-digital-service /audios-lingala-annotatees-v3.2gated audios-lingala-annotatees-v3.2 MALOBA Lingala speech corpus: 18,767 annotated audio segments (58.1 hours), with the original transcriptions and a harmonised version in the MALOBA spelling convention. MALOBA (digitisation and acceleration of Lingala) is led by Congo Digital Services (CDS SARL) in partnership with UNDP Republic of Congo and the Ministry of Posts, Telecommunications and the Digital Economy. Licence: Nwulite Obodo Open Data License 1.0 (NOODL-1.0). Access is gated.… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v3.2.audioautomatic-speech-recognition10K<n<100K0 likes50 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.