datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chatssynth-qa-taste-codec-chat
Synthetic QA Taste-S Codec Chat
18571 single-turn Traditional Chinese QA utterances with synthesized speech, 21.6 hours of
audio before codec extraction.
Assistant speech is represented as:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
Each text token is followed by its 16 Taste-S FSQ codes (codebooks a..p).
Configurations
default — messages (user question + assistant <SAY> speech), audio, and answer text.
Statistics
Utterances: 18571… See the full description on the dataset page: https://huggingface.co/datasets/yilele/synth-qa-taste-codec-chat.
