Team Ai
Datasetpublic

shangeth/librispeech-mimi-codes

LibriSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LibriSpeech corpus — multi-speaker English audiobook readings from the LibriVox project. This dataset contains codes only, not audio. For waveforms, use any of the LibriSpeech mirrors (e.g. openslr/librispeech_asr); these codes let you skip the ~hours of GPU extraction needed to train Mimi-based speech models. Schema One row per utterance: Column Type Notes id string… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/librispeech-mimi-codes.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes62downloads
12 commits on main
45202205mo ago

Add Links section (extraction code, Wren project, TTS models)

shangeth
4b004676mo ago

Update citation to Wren 2026

shangeth
83e82c36mo ago

Add train_other_500

shangeth
d2294ed6mo ago

Add train_clean_360

shangeth
1345bde6mo ago

Add train_clean_100

shangeth
b339e486mo ago

Add test_other

shangeth
3855a476mo ago

Add test_clean

shangeth
ae6db336mo ago

Add dev_other

shangeth
b1bce406mo ago

Add dev_clean

shangeth
ec620986mo ago

Add dataset card

shangeth
18b345c6mo ago

Upload LibriSpeech Mimi codes

shangeth
a4fb3e66mo ago

initial commit

shangeth