mesolitica/Zeroshot-Audio-Classification-Instructions
Zeroshot-Audio-Classification-Instructions Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label, VGGSound FSD50k Nonspeech7k urbansound8K VocalSound Emotion Gender ESD Emotion Age Language TAU Urban Acoustic Scenes 2022 CochlScene BirdCLEF_2021 EmoBox AudioSet We also converted huge WAV files into MP3 16k sample rate to reduce storage size. To prevent leakage, please do not include test set in training… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.
Zeroshot-Audio-Classification-Instructions
Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label,
- VGGSound
- FSD50k
- Nonspeech7k
- urbansound8K
- VocalSound
- Emotion
- Gender
- ESD Emotion
- Age
- Language
- TAU Urban Acoustic Scenes 2022
- CochlScene
- BirdCLEF_2021
- EmoBox
- AudioSet
We also converted huge WAV files into MP3 16k sample rate to reduce storage size.
To prevent leakage, please do not include test set in training session.
how to prepare the dataset
huggingface-cli download \
mesolitica/Zeroshot-Audio-Classification-Instructions \
--include "*.zip" \
--repo-type "dataset" \
--local-dir './'
huggingface-cli download \
mesolitica/Audio-Adversarial-Instructions \
--include "*.zip" \
--repo-type "dataset" \
--local-dir './'
huggingface-cli download \
mesolitica/Animal-Sound-Instructions \
--include "*.zip" \
--repo-type "dataset" \
--local-dir './'
huggingface-cli download \
mesolitica/EmoBox \
--include "*.zip" \
--repo-type "dataset" \
--local-dir './'
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3 unzip.pyAcknowledgement
Special thanks to https://www.sns.com.my and Nvidia for 8x H100 node!
