shangeth/librispeech-mimi-codes
LibriSpeech — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for the LibriSpeech corpus — multi-speaker English audiobook readings from the LibriVox project. This dataset contains codes only, not audio. For waveforms, use any of the LibriSpeech mirrors (e.g. openslr/librispeech_asr); these codes let you skip the ~hours of GPU extraction needed to train Mimi-based speech models. Schema One row per utterance: Column Type Notes id string… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/librispeech-mimi-codes.
Add Links section (extraction code, Wren project, TTS models)
Update citation to Wren 2026
Add train_other_500
Add train_clean_360
Add train_clean_100
Add test_other
Add test_clean
Add dev_other
Add dev_clean
Add dataset card
Upload LibriSpeech Mimi codes
initial commit
