psdn-ai/numo-indic-speech
Numo Indic Speech Dataset - Open source Numo Indic Speech is a quality-filtered, crowdsourced speech corpus collected by Poseidon AI, Inc. It contains 353,349 recordings totaling 3,845.5 hours across Bengali, Hindi, Tamil, and Telugu. The corpus consists of scripted, single-speaker read speech for automatic speech recognition (ASR), with reference transcripts, speaker metadata, and automated validation metrics for every recording. Corpus 3,845.5 hours · 353… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/numo-indic-speech.
3132
No card is published for this repository, or it could not be fetched from Hugging Face right now.
