Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /twi-health-asr-gemini-500hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Health Speech Dataset Gemini (500 hours) A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on health and wellness. Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.audioautomatic-speech-recognition10K<n<100K0 likes1.8k downloads2mo agoHugging Face02shb777 /gemini-flash-2.0-speech 🎙️ Gemini Flash 2.0 Speech Dataset This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English. 🏅 #1 Trending Audio Dataset in Feb 2025 🏅 Used in training of Kokoro TTS and LLaSA 1B 〽️ Stats Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours) Average duration: 10.83 seconds Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.audiotext-to-speech10K<n<100K60 likes959 downloads1y agoHugging Face03ghananlpcommunity /twi-health-asr-gemini-500hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Health Speech Dataset Gemini (500 hours) A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on health and wellness. Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs.audioautomatic-speech-recognition10K<n<100K1 likes717 downloads4mo agoHugging Face04ghananlpcommunity /twi-health-asr-gemini-500hrs-ipa Twi Health Speech — Audio, Transcript and IPA Twi health-domain speech with both a written transcript and an IPA phoneme sequence read off the audio by ASR. Built from ghananlpcommunity/twi-health-asr-gemini-500hrs by adding the IPA column. from datasets import load_dataset ds = load_dataset("ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa", split="train") ds[0]["audio"] # decoded waveform, 16 kHz ds[0]["transcription"] # transcript ds[0]["ipa"]… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa.audioautomatic-speech-recognition10K<n<100K0 likes270 downloads2mo agoHugging Face05Chapimenge /amharic-gemination-lexicon Amharic Gemination Lexicon v3 Which consonants are doubled in each of 86,022 Amharic words, for every reading of the word, and where each doubling comes from. Built by Dataset.ET with the HornMorpho morphological analyzer, corrected with rules that native listening, Armbruster's hand-marked verb tables and recorded speech agree on. Word types 86,022 (856,734 corpus tokens) Analyzed by HornMorpho 56,964 types, 84.0% of tokens Words with a geminate in the top… See the full description on the dataset page: https://huggingface.co/datasets/Chapimenge/amharic-gemination-lexicon.tabulartext-to-speech10K<n<100K1 likes176 downloads15d agoHugging Face06aharalambieva /target-words-geminitts Target-Word (TW) Evaluation Set Synthetic speech clips for 101 rare drug terms, intended for evaluation only — measuring how well an ASR system recognises rare / out-of-vocabulary medical vocabulary (target-word WER / CER / recall). Each clip reads a real DailyMed sentence containing one target drug name, synthesised with Google Gemini TTS across multiple voices. This is the frozen evaluation set from the master's thesis "Audio-free lexical adaptation of Whisper's decoder"… See the full description on the dataset page: https://huggingface.co/datasets/aharalambieva/target-words-geminitts.audioautomatic-speech-recognition1K<n<10K0 likes15 downloads3mo agoHugging Face07BernardoAI /wpp_pav_transcrito_gemini 🎤 Transcrições WhatsApp - Google Gemini 2.0 Flash Este dataset contém transcrições de mensagens de áudio do WhatsApp geradas usando Google Gemini 2.0 Flash. 📋 Descrição Origem: Mensagens de áudio do WhatsApp em português brasileiro Modelo: Google Gemini 2.0 Flash Preço: $0.075/1M tokens Total de amostras: 198 Formato de áudio: WAV (16kHz) Idioma: Português brasileiro Modelo de IA multimodal avançado da Google com capacidades de análise contextual de áudio.… See the full description on the dataset page: https://huggingface.co/datasets/BernardoAI/wpp_pav_transcrito_gemini.audioautomatic-speech-recognitionn<1K0 likes14 downloads1y agoHugging Face08huckiyang /gemma-4-public-bench-evalgated Gemma 4 (e4b & 12b) — Public ASR Benchmark (Decoded Hypotheses + WER) Decoded transcripts and word-level error metrics from Gemma 4 Unified (the encoder-free, natively audio-capable models) run as automatic speech recognition (ASR) systems on three standard English test sets. Two models are evaluated — gemma4:e4b (8B params) and gemma4:12b. Everything was produced locally with ollama; the evaluation tool (eval_asr.py) is included so the numbers are fully reproducible. Gemma 4… See the full description on the dataset page: https://huggingface.co/datasets/huckiyang/gemma-4-public-bench-eval.tabularautomatic-speech-recognition10K<n<100K1 likes9 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.