datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
en_spontaneous_profanityDataset from this article: Multimodal prediction of profanity based on speech analysis includes "cleaned_cv.zip" which consists of "Common voice" dataset's records with profanity speech. It can be treated as train data. "collected_records.zip" is manually collected from YouTube videos with spontaneous speech with profanity.
Annotation files have the names of files and transcribed speech.
The dataset was used to test solution for real-time profanity prediction, solution is located here: github… See the full description on the dataset page: https://huggingface.co/datasets/Lameus/en_spontaneous_profanity.profanity-speech-suroboyoan
Dataset Audio Perkataan Vulgar Bahasa Jawa Dialek Surabaya
Deskripsi
Dataset ini berisi audio rekaman percakapan dalam bahasa Jawa dialek Surabaya yang mengandung perkataan vulgar. Setiap rekaman dilengkapi dengan transkripsi teks yang sesuai.Dataset ini dibuat sebagai bagian dari penelitian skripsi saya dengan tujuan untuk mendukung analisis dan pengembangan dalam bidang deteksi perkataan vulgar dalam bahasa Jawa dialek Surabaya menggunakan teknologi speech-to-text… See the full description on the dataset page: https://huggingface.co/datasets/Jaal047/profanity-speech-suroboyoan.ProfanityDetection_TAPAD
