Team Ai
Datasetpublic

shangeth/mls-mimi-codes

Multilingual LibriSpeech (MLS) — Mimi Codes Pre-extracted Kyutai Mimi neural-codec tokens for Multilingual LibriSpeech — LibriVox audiobooks in 7 non-English languages. English is intentionally excluded. For English Mimi codes, use: shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits) shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native) shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents) shangeth/jenny-mimi-codes —… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes107downloads
Dataset Card

Multilingual LibriSpeech (MLS) — Mimi Codes

Pre-extracted Kyutai Mimi neural-codec tokens for Multilingual LibriSpeech — LibriVox audiobooks in 7 non-English languages.

English is intentionally excluded. For English Mimi codes, use:

Configs (languages)

One HF dataset config per language:

ConfigLanguageISOApprox hours (train)
dutchDutchnl~1.5k
frenchFrenchfr~1.1k
germanGermande~3.3k
italianItalianit~250
polishPolishpl~100
portuguesePortuguesept~160
spanishSpanishes~920

Splits (per config)

SplitDescription
trainfull training set
devdevelopment
testtest
9_hourslow-resource ~9h training subset
1_hourslow-resource ~1h training subset

Merged all config

A virtual all config aliases the per-language parquets via glob — no extra storage, just a YAML entry. Use it to train a single multilingual model:

python
ds = load_dataset("shangeth/mls-mimi-codes", "all", split="train")

Caveat: there is no language column, and speaker_id is per-language — so speaker IDs may collide across languages (e.g. German speaker 12345 is unrelated to French speaker 12345). If you need clean per-language speaker bookkeeping, load each language config separately.

Schema

ColumnTypeNotes
idstringutterance ID, format {speaker}_{chapter}_{segment}
textstringtranscript, mixed-case as-is from MLS
speaker_idint32speaker ID (parsed from MLS string)
chapter_idint32chapter ID
codesint16[k=8][n_frames]Mimi codebook indices @ 12.5 fps
n_framesint32
k_codebooksint328

Extraction details

Usage

python
from datasets import load_dataset
import torch

ds = load_dataset("shangeth/mls-mimi-codes", "german", split="dev")
ex = ds[0]
codes = torch.tensor(ex["codes"], dtype=torch.long)  # [8, n_frames]
print(ex["id"], "| speaker:", ex["speaker_id"], "|", ex["text"][:60])

# Decode back to 24 kHz audio
from transformers import MimiModel
mimi = MimiModel.from_pretrained("kyutai/mimi").cuda().eval()
with torch.no_grad():
    wav = mimi.decode(codes.unsqueeze(0).cuda()).audio_values[0].cpu()

Links

Citation

bibtex
@misc{wren2026,
  title  = {Wren: A Family of Small Open-Weight Models for Unified Speech-Text Modelling},
  author = {Shangeth Rajaa},
  year   = {2026},
  url    = {https://github.com/shangeth/wren}
}

@inproceedings{pratap2020mls,
  title     = {MLS: A Large-Scale Multilingual Dataset for Speech Research},
  author    = {Pratap, Vineel and Xu, Qiantong and Sriram, Anuroop and Synnaeve, Gabriel and Collobert, Ronan},
  booktitle = {Interspeech},
  year      = {2020}
}

License

CC-BY-4.0 (inherited from MLS / LibriVox).