jerpint/vox-cloned-data
CommonVoice Clones This dataset consists of recordings taken from the CommonVoice english dataset. Each voice and transcript are used as input to a voice cloner, and generate a cloned version of the voice and text. TTS Models We use the following high-scoring models from the TTS leaderboard: playHT metavoice StyleTTSv2 XttsV2 Model Comparisons To facilitate data exploration, check out this HF space 🤗, which allows you to listen to all clones… See the full description on the dataset page: https://huggingface.co/datasets/jerpint/vox-cloned-data.
CommonVoice Clones
This dataset consists of recordings taken from the CommonVoice english dataset. Each voice and transcript are used as input to a voice cloner, and generate a cloned version of the voice and text.
TTS Models
We use the following high-scoring models from the TTS leaderboard:
- playHT
- metavoice
- StyleTTSv2
- XttsV2
Model Comparisons
To facilitate data exploration, check out this HF space 🤗, which allows you to listen to all clones from a given source (e.g. commonvoice).
Data Sampling
The subset of the data was chosen such that a balanced amount of male/female voices are included. We also include a suggested train, validation and test split. Unique voice IDs do not overlap between the splits, i.e. a voice used to clone a sample in train will never be found in validation and test.
