AudioCC-Lab/PICSAFEv1
Speech Quality Test Labels PICSAFEv1 is a multi-source annotated test dataset for evaluating speech quality assessment and audio data filtering methods. It contains 10,728 audio samples drawn from 14 source datasets, with annotations from a vocabulary of 33 tags. These tags describe recording provenance, speech styles, speaking rate and pitch, speaker attributes, noise, reverberation, distortion, and transcript errors. These annotations support benchmarking quality metrics and… See the full description on the dataset page: https://huggingface.co/datasets/AudioCC-Lab/PICSAFEv1.
Speech Quality Test Labels
PICSAFEv1 is a multi-source annotated test dataset for evaluating speech quality assessment and audio data filtering methods. It contains 10,728 audio samples drawn from 14 source datasets, with annotations from a vocabulary of 33 tags. These tags describe recording provenance, speech styles, speaking rate and pitch, speaker attributes, noise, reverberation, distortion, and transcript errors. These annotations support benchmarking quality metrics and filtering methods, selecting thresholds, analyzing errors across audio conditions, and studying data selection for automatic speech recognition and text-to-speech using independently acquired source audio.
The dataset provides a shared evaluation set for studying which audio samples a filtering system accepts or rejects, and why. Its tag annotations support both detailed analysis of individual audio characteristics and binary evaluation under the strict, mid, and loose filtering policies described below. Acceptance is policy-dependent: a negative label indicates exclusion under the selected policy, rather than that the recording is unsuitable for every task.
When using PICSAFEv1 for threshold selection or model development, keep a separate held-out subset for final evaluation to avoid reporting performance on the same samples used for tuning.
It publishes this README, `metadata.jsonl`, and a tag-to-binary conversion script, because the 14 source datasets have different copyright and access terms.
The manifest contains the 10,728-sample PICSAFEv1 subset. It is not a license to redistribute the underlying recordings.
Metadata fields
file_name: audio basename.id: globally unique sample identifier with the[source]prefix.source: source dataset name.source_path: path relative to the origin directory.tags: list of labels.
Source datasets and official download paths
Label schema
In article, it is described that the 33 labels are:
Nspk, distortion, emotional, enhanced, excessive_sibilance, fast_speed, female, high_pitch, imperceptible_noise_level, instantaneous_noise, low_noise_level, low_pitch, male, microphone_popping, mid_high_noise_level, music_or_effect, nb_noise, non_binary, non_speech, read_speech, real_recording, regular_pitch, regular_speed, reverberation, singing_voice, slow_speed, speech_overlap, spontaneous_speech, synthetic, vocal_sound, wb_noise, whispered_speech, and wrong_transcript.
Converting tags to binary labels
Labels are derived from the human tags, not predictions from a decision tree or metric thresholds.
1: accepted under the selected policy (positive).0: rejected because at least one negative tag is present (negative).
The label is 0 if any annotation tag belongs to the mode-specific negative set below or the additional --negative_tags set; otherwise it is 1.
mid allows Nspk and low_noise_level; loose additionally allows microphone_popping and excessive_sibilance. Other tags do not independently cause rejection. In particular, non_speech, wrong_transcript, singing_voice, vocal_sound, nb_noise, wb_noise, and instantaneous_noise are not default negative tags. Label 1 means policy acceptance, not a universal claim of clean speech or a correct transcript. Add rejection tags explicitly if needed. Names are case-sensitive (Nspk). An empty tag list yields 1, matching the reference function, but does not establish annotation completeness. Missing tag fields and unknown tag names are errors.
Run from the repository directory:
# Default policy.
python3 tags_to_binary.py --input metadata.jsonl --output /tmp/picsafe_strict.jsonl
# Alternative policy (also supports loose).
python3 tags_to_binary.py --input metadata.jsonl --output /tmp/picsafe_mid.jsonl --positive_mode mid
# Custom policy, different from the default labels.
python3 tags_to_binary.py --input metadata.jsonl --output /tmp/picsafe_custom.jsonl --negative_tags non_speech wrong_transcript
# Original annotations: UID and semicolon-separated Tags columns.
python3 tags_to_binary.py --input annotations.tsv --output /tmp/picsafe_from_tsv.jsonlOutput JSONL preserves all original fields and adds binary_label and label_policy (the selected positive_mode and additional negative_tags). TSV input also adds id from UID and tags from Tags. The script prints class counts, validates unique IDs and tag names, and refuses to overwrite existing output files. The source metadata.jsonl is unchanged. Join with metric scores using globally unique id; basenames can collide. To reproduce a threshold-search run, use the same mode, additional negative tags, and sample subset.
Licensing and limitations
This repository does not grant, aggregate, or override any source-dataset license. Obtain the audio directly from the source providers and follow their current terms, attribution requirements, access restrictions, and applicable privacy/publicity rules. Some sources require non-commercial research use, registration, a license, or direct permission from the authors.
