Team Ai
Datasetpublic

shangeth/jenny-mimi-codes

Jenny TTS — Mimi Codes Pre-extracted Kyutai Mimi tokens for the Jenny TTS Dataset — a single female speaker, ~30h, clean studio-quality recordings. Apache-2.0 license. Schema Column Type Notes id string e.g. jenny_0 text string mixed-case with punctuation codes int16[k=8][n_frames] Mimi codebook indices @ 12.5 fps n_frames int32 k_codebooks int32 8 No speaker_id column — single speaker dataset. Extraction details Source:… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/jenny-mimi-codes.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes138downloads
Dataset Card

Jenny TTS — Mimi Codes

Pre-extracted Kyutai Mimi tokens for the Jenny TTS Dataset — a single female speaker, ~30h, clean studio-quality recordings. Apache-2.0 license.

Schema

ColumnTypeNotes
idstringe.g. jenny_0
textstringmixed-case with punctuation
codesint16[k=8][n_frames]Mimi codebook indices @ 12.5 fps
n_framesint32
k_codebooksint328

No speaker_id column — single speaker dataset.

Extraction details

Usage

python
from datasets import load_dataset
import torch

ds = load_dataset("shangeth/jenny-mimi-codes", split="train")
ex = ds[0]
codes = torch.tensor(ex["codes"], dtype=torch.long)  # [8, n_frames]
print(ex["text"])

Links

Citation

bibtex
@misc{wren2026,
  title  = {Wren: A Family of Small Open-Weight Models for Unified Speech-Text Modelling},
  author = {Shangeth Rajaa},
  year   = {2026},
  url    = {https://github.com/shangeth/wren}
}

License

Apache-2.0.