Team Ai
Datasetpublic

espnet/yodas3

YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data. For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.

sourceHugging Facecc-by-3.0updated 5d agoView on Hugging Face
206likes128kdownloads

espnet/yodas3 · main · files are served by the source, never re-hosted here