Team Ai
Datasetpublic

shangeth/libritts-r-mimi-codes

LibriTTS-R — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for LibriTTS-R — a speech-restored version of LibriTTS built specifically for TTS research. Why LibriTTS-R instead of LibriSpeech? LibriSpeech LibriTTS-R Purpose ASR TTS Sample rate 16 kHz 24 kHz (Mimi-native, no resampling) Segmentation Arbitrary chunks Sentence-level Punctuation Stripped (ALL CAPS) Preserved Audio quality Raw amateur Speech restoration applied No resampling is needed… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/libritts-r-mimi-codes.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes71downloads
Dataset Card

LibriTTS-R — Mimi Codes

Pre-extracted Kyutai Mimi neural-codec tokens for LibriTTS-R — a speech-restored version of LibriTTS built specifically for TTS research.

Why LibriTTS-R instead of LibriSpeech?

LibriSpeechLibriTTS-R
PurposeASRTTS
Sample rate16 kHz24 kHz (Mimi-native, no resampling)
SegmentationArbitrary chunksSentence-level
PunctuationStripped (ALL CAPS)Preserved
Audio qualityRaw amateurSpeech restoration applied

No resampling is needed — 24 kHz matches Mimi exactly.

Schema

ColumnTypeNotes
idstringe.g. 84_121123_000003_000000
textstringnormalized text, mixed-case with punctuation preserved
speaker_idint32LibriTTS speaker ID
codesint16[k=8][n_frames]Mimi codebook indices @ 12.5 fps
n_framesint32
k_codebooksint328

Extraction details

  • —Codec: `kyutai/mimi` @ 24 kHz, 12.5 fps
  • —Codebooks: all 8 extracted. Slice codes[:k] for fewer.
  • —Source: OpenSLR 141

Splits

HF SplitSource~Rows
train_clean_100train-clean-100~33.2k
train_clean_360train-clean-360~116k
train_other_500train-other-500~205k
dev_cleandev-clean~2.7k
dev_otherdev-other~2.9k
test_cleantest-clean~2.6k
test_othertest-other~2.9k

Usage

python
from datasets import load_dataset
import torch

ds = load_dataset("shangeth/libritts-r-mimi-codes", split="train_clean_100")
ex = ds[0]
codes = torch.tensor(ex["codes"], dtype=torch.long)  # [8, n_frames]
print(ex["text"])  # "He hoped there would be stew for dinner, turnips and carrots."

Links

Citation

bibtex
@misc{wren2026,
  title  = {Wren: A Family of Small Open-Weight Models for Unified Speech-Text Modelling},
  author = {Shangeth Rajaa},
  year   = {2026},
  url    = {https://github.com/shangeth/wren}
}

@inproceedings{koizumi2023libritts,
  title     = {LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus},
  author    = {Koizumi, Yuma and others},
  booktitle = {Interspeech},
  year      = {2023}
}

License

CC-BY-4.0 (inherited from LibriTTS-R).