Team Ai
Datasetpublicgated

psdn-ai/numo-indic-speech

Numo Indic Speech Dataset - Open source Numo Indic Speech is a quality-filtered, crowdsourced speech corpus collected by Poseidon AI, Inc. It contains 353,349 recordings totaling 3,845.5 hours across Bengali, Hindi, Tamil, and Telugu. The corpus consists of scripted, single-speaker read speech for automatic speech recognition (ASR), with reference transcripts, speaker metadata, and automated validation metrics for every recording. Corpus 3,845.5 hours · 353… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/numo-indic-speech.

sourceHugging Facecc-by-4.0updated 8d agoView on Hugging Face
3likes132downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.