datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-eroge-voice-rag
Japanese-Eroge-Voice-V2
Description
This is the successor to the Japanese-Eroge-Voice dataset. It consists of a significantly larger collection of audio-transcription pairs extracted from Japanese eroge (adult games).
Note on Versioning: There is no overlap between this dataset (V2) and the previous version. All audio clips and transcriptions in V2 are distinct from those in the original version, providing entirely new data for research.
This version (V2) expands… See the full description on the dataset page: https://huggingface.co/datasets/masapon05/japanese-eroge-voice-rag.bangla_speech_corpus_v1
Dataset Card for Bangla Speech Corpus v1.0 (bangla_speech_corpus_v1.0)
Mandatory Human Audit & Quality Disclaimer
"This corpus is an automatically quality-filtered Bangla speech corpus with a stratified 200-sample human audit."
"Only 200 samples were manually audited. The remaining samples were not manually verified."
This dataset is not "fully human verified", "completely manually validated", or "100% human reviewed". Quality guarantees rest upon rigorous… See the full description on the dataset page: https://huggingface.co/datasets/ragibhasan/bangla_speech_corpus_v1.TamilVoiceCorpus
Tamil Conversational ASR Dataset
This is a dataset for Automatic Speech Recognition (ASR) focused on conversational Tamil, collected from various public sources on the web. Each sample is a short audio clip (averaging 10 seconds) paired with its corresponding transcription.
Dataset Summary
Language: Tamil (ta)
Domain: Conversational speech
Average Duration per Clip: ~10 seconds
Format: Audio (.wav) + text
Sample Rate: 16kHz recommended
Total Examples:
PureVox:
Train: 12… See the full description on the dataset page: https://huggingface.co/datasets/ragunath-ravi/TamilVoiceCorpus.embedded_world_2026_rag_tts
