pupengleileileilei/CFAD
CFAD Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test set (arXiv 2207.12308), for speech anti-spoofing and synthetic / deepfake voice detection on Mandarin Chinese speech. Overview CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean version's two test partitions: test_seen — spoof systems and real corpora also present in the train/dev splits. test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/pupengleileileilei/CFAD.
CFAD
Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test set (arXiv 2207.12308), for speech anti-spoofing and synthetic / deepfake voice detection on Mandarin Chinese speech.
Overview
CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean version's two test partitions:
- `test_seen` — spoof systems and real corpora also present in the train/dev splits.
- `test_unseen` — spoof systems and real corpora held out of training, to measure generalisation to unseen attacks.
The label is binary: bonafide (genuine human speech) vs. spoof (synthesized / vocoded speech). It is derived from the source directory layout — clips under fake_clean/ are spoof; clips under real_clean/ are bonafide. There is no separate protocol file.
One aishell1 outlier (BAC009S0764W0123.wav, a 178 MB / ~97 min concatenation artifact) is dropped, so the bonafide count is 20,999 rather than 21,000.
License & redistribution
CFAD is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, which permits redistribution with attribution — see LICENSE.txt. Audio is re-encoded to 16 kHz mono FLAC for a uniform schema; labels are unmodified.
Schema
notes example:
{"utterance_id": "CFAD_test_seen_clean_fake_clean_gl_SSB07800001_gl", "partition": "test_seen", "source": "gl"}Quick Start
from datasets import load_dataset
ds = load_dataset("SpeechAntiSpoofingBenchmarks/CFAD", split="test")
print(ds[0])Stats
Source provenance
- Paper: https://arxiv.org/abs/2207.12308
- Scope: the clean version's
test_seen_clean+test_unseen_cleanpartitions (thecodec/noisyrobustness variants and train/dev splits are not included). - Labels derived from the source directory layout (
fake_clean/= spoof,real_clean/= bonafide).
Evaluation
For evaluation instructions and submission format, see `submissions/README.md`.
Citation
Paper: Haoxin Ma et al., CFAD: A Chinese Dataset for Fake Audio Detection, arXiv:2207.12308 (2022).
@article{ma2022cfad,
title = {CFAD: A Chinese Dataset for Fake Audio Detection},
author = {Ma, Haoxin and Yi, Jiangyan and Wang, Chenglong and Yan, Xinrui and
Tao, Jianhua and Wang, Tao and Wang, Shiming and Fu, Ruibo},
journal = {arXiv preprint arXiv:2207.12308},
year = {2022}
}Maintainer
Maintained by Kirill Borodin (SpeechAntiSpoofingBenchmarks).
- Email: ~~k.n.borodin@mtuci.ru~~ (deprecated — use kborodin.research@gmail.com)
- Telegram: @korallll_ai
