shangeth/libritts-r-mimi-codes
LibriTTS-R — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for LibriTTS-R — a speech-restored version of LibriTTS built specifically for TTS research. Why LibriTTS-R instead of LibriSpeech? LibriSpeech LibriTTS-R Purpose ASR TTS Sample rate 16 kHz 24 kHz (Mimi-native, no resampling) Segmentation Arbitrary chunks Sentence-level Punctuation Stripped (ALL CAPS) Preserved Audio quality Raw amateur Speech restoration applied No resampling is needed… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/libritts-r-mimi-codes.
LibriTTS-R — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens for LibriTTS-R — a speech-restored version of LibriTTS built specifically for TTS research.
Why LibriTTS-R instead of LibriSpeech?
No resampling is needed — 24 kHz matches Mimi exactly.
Schema
Extraction details
- Codec: `kyutai/mimi` @ 24 kHz, 12.5 fps
- Codebooks: all 8 extracted. Slice
codes[:k]for fewer. - Source: OpenSLR 141
Splits
Usage
from datasets import load_dataset
import torch
ds = load_dataset("shangeth/libritts-r-mimi-codes", split="train_clean_100")
ex = ds[0]
codes = torch.tensor(ex["codes"], dtype=torch.long) # [8, n_frames]
print(ex["text"]) # "He hoped there would be stew for dinner, turnips and carrots."Links
- Dataset extraction code: github.com/shangeth/wren-datasets
- Wren research project: github.com/shangeth/wren
- TTS models trained on these codes: github.com/shangeth/wren-tts
Citation
@misc{wren2026,
title = {Wren: A Family of Small Open-Weight Models for Unified Speech-Text Modelling},
author = {Shangeth Rajaa},
year = {2026},
url = {https://github.com/shangeth/wren}
}
@inproceedings{koizumi2023libritts,
title = {LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus},
author = {Koizumi, Yuma and others},
booktitle = {Interspeech},
year = {2023}
}License
CC-BY-4.0 (inherited from LibriTTS-R).
