RootAccess4Life/cond-id-samples
cond-ID โ audio samples Synthesized audio backing the paper "cond-ID: Conditioning-Space Identity Redirection for Speaker Unlearning in Zero-Shot TTS." ~2,450 clips: every backbone, every baseline, and the relearn stress test. ๐ Prefer listening in the browser? โ Live demo Space The one thing to listen for Each speaker appears as a pair: clip meaning cB_โฆ baseline clone โ the un-edited model cloning that speaker. This is the voice being copied. cE_โฆโฆ See the full description on the dataset page: https://huggingface.co/datasets/RootAccess4Life/cond-id-samples.
cond-ID โ audio samples
Synthesized audio backing the paper "cond-ID: Conditioning-Space Identity Redirection for Speaker Unlearning in Zero-Shot TTS." ~2,450 clips: every backbone, every baseline, and the relearn stress test.
๐ Prefer listening in the browser? โ Live demo Space
The one thing to listen for
Each speaker appears as a pair:
Play cB then cE for the same speaker: intelligible speech in both, but the voice identity should no longer match. cRL asks whether an adversary can undo that.
Filename convention
{cB|cE|cRL}_{speaker}_{sentence}.wavSpeaker ids are VCTK (p225, p272, โฆ) or LibriTTS (103, 118, โฆ) depending on the split; {sentence} indexes a fixed sentence bank, so the same index is the same text across every directory. TruS-F5 clips carry an extra seed field (cB_s0_p225_0.wav).
Directory map
Main result โ cond-ID on three backbones
Stress tests
Baselines (each in {method}_{backbone}_clones/)
additive_gaussian ยท gradient_ascent_to_forget ยท random_reference_substitution ยท task_arithmetic ยท zero_pooled_identity
TruS on F5-TTS
Loading
Stream a single clip without cloning the whole set:
from huggingface_hub import hf_hub_download
p = hf_hub_download(
repo_id="RootAccess4Life/cond-id-samples",
filename="xtts_clones/cE_103_0.wav",
repo_type="dataset",
)Or grab one backbone:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="RootAccess4Life/cond-id-samples",
repo_type="dataset",
allow_patterns=["xtts_clones/*"],
)Provenance and licensing
Clips are synthesized by XTTS-v2, Tortoise-TTS, IndexTTS-1.5 and F5-TTS from reference audio in CSTR VCTK-Corpus-0.92 (CC-BY-4.0, DOI 10.7488/ds/2645) and LibriTTS (CC-BY-4.0, OpenSLR SLR60). This dataset is released CC-BY-4.0 to match those sources.
The generating models keep their own licenses โ notably XTTS-v2 is non-commercial (CPML). Those terms constrain the models, not these audio files, but review them before building on the pipeline.
Intended use and limits
Released for research and demonstration: verifying the paper's claims by ear, and supporting work on speaker unlearning, privacy, and the right to be forgotten in speech synthesis.
These are synthetic clones of real speakers, produced to demonstrate removing voice identity. Do not use them to impersonate anyone, to train or evaluate systems that attribute speech to the original speakers, or as evidence any of these people said these words โ they did not. The cB_ clips in particular are deliberate voice clones and should be handled accordingly.
Related
- ๐ป Code: https://github.com/pujariaditya/cond-id-tts-unlearning
- ๐ Live demo: https://huggingface.co/spaces/RootAccess4Life/cond-id-demo
- ๐ฆ Mirrors: XTTS-v2 ยท IndexTTS-1.5
Citation
@misc{pujari_condid,
title = {cond-ID: Conditioning-Space Identity Redirection for Speaker Unlearning in Zero-Shot TTS},
author = {Pujari, Aditya and Rattani, Ajita},
note = {Preprint},
}