AudioCC-Lab/PICSAFEv1
Speech Quality Test Labels PICSAFEv1 is a multi-source annotated test dataset for evaluating speech quality assessment and audio data filtering methods. It contains 10,728 audio samples drawn from 14 source datasets, with annotations from a vocabulary of 33 tags. These tags describe recording provenance, speech styles, speaking rate and pitch, speaker attributes, noise, reverberation, distortion, and transcript errors. These annotations support benchmarking quality metrics and… See the full description on the dataset page: https://huggingface.co/datasets/AudioCC-Lab/PICSAFEv1.
Introduce PICSAFEv1 purpose and applications
Document binary tag policies and add metadata conversion script
Document tag to binary label conversion
Update dataset README
Add metadata manifest and source download links
Remove audio data and document source downloads
Fix dataset card data files config
Add PICSAFEv1 audio dataset
initial commit
