Team Ai
20 results

Multilingual-TTS

multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.7k downloads4mo agoHugging Facenineninesix /multilingual-tts-benchmark Multilingual Speech Benchmark for Zero-Shot TTS A voice-cloning and intelligibility benchmark for 8 language subsets, with 10,100 examples selected from Common Voice 17.0. Each example supplies a speaker reference and an independently selected target text, with human recordings as WER/CER and speaker-similarity anchors when the corresponding audio is available. Corpus WER and the WavLM-FT evaluation follow seed-tts-eval. Version 3.1. Adds ru, kk through the same S1–S4 selection… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.audiotext-to-speech100K<n<1M0 likes1.8k downloads4d agoHugging Facemalaysia-ai /Multilingual-TTS Multilingual-TTS A large multilingual corpus for pretraining TTS/STT models, gathered and normalized from 230+ public sources. ~191k hours of audio across 150+ languages, organized into 1,544 dataset configs and tokenized to 34.5B NeuCodec speech tokens (50 Hz, single-codebook) over 111.1M clips. Each config is one source dataset normalized to rows of {audio_filename, text, speaker}: audio_filename — clip path inside that config's <config>_audio.zip (mono MP3). text —… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS.24 likes1.2k downloads2mo agoHugging FaceMiniMaxAI /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K47 likes1.1k downloads1y agoHugging FaceCong123779 /matcha-tts-multilingual-datasets 🎙️ Matcha-TTS Multilingual Master Speech & Acoustic Datasets Kho dữ liệu âm thanh đa ngôn ngữ (Tiếng Việt & Tiếng Anh) chuẩn phòng thu SOTA, được tiền xử lý và thẩm định toàn diện (SQUIM noise assessment, trích xuất vector âm sắc ECAPA-TDNN 192-d, căn chỉnh phụ đề JSON/timestamps, Mel-spectrogram & F0 pitch contours) phục vụ tối ưu cho việc huấn luyện và suy luận mô hình Matcha-TTS kết hợp Vocoder Vocos. 📌 I. Tổng Quan Các Bộ Dữ Liệu Trên Hugging Face Toàn bộ… See the full description on the dataset page: https://huggingface.co/datasets/Cong123779/matcha-tts-multilingual-datasets.text-to-speech100K<n<1M0 likes1.1k downloads8d agoHugging Faceakuzdeuov /qwen3-tts-multilingual-emotional-speechaudio1M<n<10M0 likes962 downloads28d agoHugging Face