Team Ai
Datasetpublic

shangeth/librispeech-mimi-codes

LibriSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LibriSpeech corpus — multi-speaker English audiobook readings from the LibriVox project. This dataset contains codes only, not audio. For waveforms, use any of the LibriSpeech mirrors (e.g. openslr/librispeech_asr); these codes let you skip the ~hours of GPU extraction needed to train Mimi-based speech models. Schema One row per utterance: Column Type Notes id string… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/librispeech-mimi-codes.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes62downloads
../
filedev_clean-00000-of-00001.parquet2.8 MBdownload
filedev_other-00000-of-00001.parquet2.7 MBdownload
filetest_clean-00000-of-00001.parquet2.8 MBdownload
filetest_other-00000-of-00001.parquet2.8 MBdownload
filetrain_clean_100-00000-of-00001.parquet51.1 MBdownload
filetrain_clean_360-00000-of-00001.parquet184.6 MBdownload
filetrain_other_500-00000-of-00001.parquet252.0 MBdownload

shangeth/librispeech-mimi-codes · main · files are served by the source, never re-hosted here