Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01geronimobasso /drone-audio-detection-samples Dataset Description Drone Audio Detection Samples (DADS) is currently the largest publicly available drone audio database, specifically designed for developing drone detection systems using deep learning techniques. All audio files are standardized to a sample rate of 16,000 Hz, 16-bit depth, mono-channel, and vary in length from 500 milliseconds to several minutes. Most drone audio files were manually trimmed to ensure that a drone was always present in the recording. However, some… See the full description on the dataset page: https://huggingface.co/datasets/geronimobasso/drone-audio-detection-samples.audioaudio-classification100K<n<1M39 likes1.7k downloads2y agoHugging Face02garystafford /deepfake-audio-detection Deepfake Audio Detection Dataset (v4) Dataset Description This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio. What's New in v4 52% larger: Increased from 1,224 to 1,866 samples (642 new samples) Expanded TTS coverage: Added Hume AI as 6th synthetic voice… See the full description on the dataset page: https://huggingface.co/datasets/garystafford/deepfake-audio-detection.audioaudio-classification1K<n<10K6 likes1.5k downloads10mo agoHugging Face03CSALT /deepfake_detection_dataset_urdu Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset This repository contains the Urdu Deepfake Audio Dataset introduced in the ACL 2024 paper "Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset". The dataset focuses on two spoofing attacks – Tacotron and VITS TTS – and includes bonafide audio samples for comparison. The dataset construction ensures phonemic cover and balance, making it suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/CSALT/deepfake_detection_dataset_urdu.audio1K<n<10K6 likes656 downloads2y agoHugging Face04TartarusXXX /mixed-language-detection-pilot-fleurs-voices Mixed-Language Speech Detection Pilot — Native FLEURS Voices This is the native-reference revision of a 6,000-clip binary audio-classification pilot. label = 0 denotes one intended language and label = 1 denotes more than one intended language. The covered languages are Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and English (eng). What changed in this revision Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.audioaudio-classification1K<n<10K0 likes598 downloads2mo agoHugging Face05schiffman /drone-audio-detection-samples Dataset Description Drone Audio Detection Samples (DADS) is currently the largest publicly available drone audio database, specifically designed for developing drone detection systems using deep learning techniques. All audio files are standardized to a sample rate of 16,000 Hz, 16-bit depth, mono-channel, and vary in length from 500 milliseconds to several minutes. Most drone audio files were manually trimmed to ensure that a drone was always present in the recording. However… See the full description on the dataset page: https://huggingface.co/datasets/schiffman/drone-audio-detection-samples.audioaudio-classification100K<n<1M0 likes426 downloads2mo agoHugging Face06dolphinteam /OpenWhistle-Detection-Finetuning OpenWhistle Detection Finetuning Expert-annotated whistle-type detection dataset for OpenWhistle. Each example is a fixed-length 0.5 s audio window labeled with the whistle types present in that window; background/no-whistle windows are represented by an all-zero target vector. Overview Task: multi-label whistle-type detection on fixed-length audio windows Target vector: label, with one binary decision per whistle type in the order SW_Neo, SW_Luna, SW_Nikita… See the full description on the dataset page: https://huggingface.co/datasets/dolphinteam/OpenWhistle-Detection-Finetuning.audio1K<n<10K1 likes349 downloads12d agoHugging Face07laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1audio100M<n<1B4 likes319 downloads1y agoHugging Face08dhlee3000 /LMD-AI-Detection LMD AI-Generated Music Detection Benchmark (Note: The corresponding research paper will be released later.) Dataset Description The rapid advancement of AI music generation has raised growing concerns about the authenticity of digital music. While deepfake detection has been extensively studied in the audio domain, symbolic music (MIDI) remains largely unexplored. This dataset presents a comprehensive benchmark for AI-generated symbolic music detection, examining… See the full description on the dataset page: https://huggingface.co/datasets/dhlee3000/LMD-AI-Detection.audioaudio-classification10K<n<100K2 likes317 downloads27d agoHugging Face09kantu9 /vessel-detection-datasetaudio1K<n<10K1 likes256 downloads4mo agoHugging Face10ArlingtonCL2 /Barkopedia-Dog-Vocal-Detection 🐾 Dog Vocal Detection This dataset is curated from internet videos to support research in dog vocalization detection using both weak and strong supervision. It contains approximately 7,500 seconds of strongly labeled training audio Over 9,000 seconds of weakly labeled clips sourced from AudioSet are included. The dataset also provides 24 hours of unlabeled audio clips from our own collection. To simulate realistic conditions, some clips feature dogs present without barking… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia-Dog-Vocal-Detection.audion<1K3 likes255 downloads1y agoHugging Face11Abdelkareem /arabic_commands_detection Dataset Card for "arabic_commands_detection" More Information needed audio10K<n<100K0 likes227 downloads3y agoHugging Face12koyyalamudiraghavendra /deepfake-audio-detection Deepfake Audio Detection Dataset (v4) Dataset Description This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio. What's New in v4 52% larger: Increased from 1,224 to 1,866 samples (642 new samples) Expanded TTS coverage: Added Hume AI as 6th synthetic… See the full description on the dataset page: https://huggingface.co/datasets/koyyalamudiraghavendra/deepfake-audio-detection.audioaudio-classification1K<n<10K0 likes212 downloads1mo agoHugging Face13mazesmazes /turn-end-detection Turn-end detection from real ASR prefixes, with audio 100,348 labelled end-of-turn decision points over 53,140 synthesized customer-service utterances, each one paired with the 16 kHz audio it was cut from, plus the endpointing decisions 6 commercial endpointer configurations made on the same audio. The question each row poses is the one a voice agent has to answer continuously: given everything heard so far, has the caller finished speaking? Ending the turn too early talks over… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/turn-end-detection.audioaudio-classification100K<n<1M0 likes208 downloads24d agoHugging Face14amine-maazizi /tta-detection-2kaudioaudio-classification1K<n<10K0 likes202 downloads7mo agoHugging Face15Hemg /Audio-based-Voilence-detection-Datasetaudion<1K0 likes171 downloads3y agoHugging Face16RapidOrc121 /audio-emotion-detection-dataset Audio Emotion Detection Dataset Github: Audio Emotion Detection Dataset Connect with me : Linkedin Speech clips in English and Hindi annotated with emotion labels and ASR transcripts. Audio is sourced from public YouTube videos and trimmed to approximately 60 seconds per clip. Noise reduction is applied via noisereduce and silero-vad. Emotions (5 classes) Label Description angry Aggressive, confrontational speech calm… See the full description on the dataset page: https://huggingface.co/datasets/RapidOrc121/audio-emotion-detection-dataset.audioaudio-classification1 likes162 downloads4mo agoHugging Face17kuross /dl-proj-detectionaudio1K<n<10K0 likes145 downloads10mo agoHugging Face18AxonData /footstep-detection-dataset Footstep Detection Dataset — 50 Hours of Real Footstep Audio 50 hours of real footstep audio recordings for training footstep detection, sound event detection, and audio classification models. 166 manually verified files captured in natural indoor and outdoor conditions, with per-file metadata on surface, footwear, location, and background noise Contact us and share your feedback — receive additional samples for free! 😊 Key Highlights 50 hours of… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/footstep-detection-dataset.audioaudio-classificationn<1K0 likes132 downloads5mo agoHugging Face19kotangalechaitali9007 /audio-emotion-detection-dataset Audio Emotion Detection Dataset Github: Audio Emotion Detection Dataset Connect with me : Linkedin Speech clips in English and Hindi annotated with emotion labels and ASR transcripts. Audio is sourced from public YouTube videos and trimmed to approximately 60 seconds per clip. Noise reduction is applied via noisereduce and silero-vad. Emotions (5 classes) Label Description angry Aggressive, confrontational speech calm… See the full description on the dataset page: https://huggingface.co/datasets/kotangalechaitali9007/audio-emotion-detection-dataset.audioaudio-classificationn<1K0 likes109 downloads2mo agoHugging Face20duongve /Dataset_audio_threat_detectionaudio100K<n<1M2 likes103 downloads1y agoHugging Face21hoyyu1 /infant-cry-detection Dataset Documentation / 数据集说明 Introduction / 简介 This dataset is designed for infant cry detection in noisy household environments. It aims to provide a diverse, representative, and robust collection of audio samples for training and evaluating machine learning models, particularly deep neural networks. 本数据集专为嘈杂家庭环境下的婴儿哭声检测而设计。它旨在提供一个多样化、具代表性且具备鲁棒性的音频样本集合,用于训练和评估机器学习模型(尤其是深度神经网络)。 The experimental dataset was constructed by integrating multiple public datasets to… See the full description on the dataset page: https://huggingface.co/datasets/hoyyu1/infant-cry-detection.audioaudio-classification10K<n<100K2 likes96 downloads5mo agoHugging Face22Tanishq125 /deepfake-audio-detection Deepfake Audio Detection Dataset (v4) Dataset Description This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio. What's New in v4 52% larger: Increased from 1,224 to 1,866 samples (642 new samples) Expanded TTS coverage: Added Hume AI as 6th synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Tanishq125/deepfake-audio-detection.audioaudio-classification1K<n<10K0 likes92 downloads29d agoHugging Face23AxonData /infant-cry-detection-dataset Infant Cry Detection Dataset — 50+ Hours of Real Baby Cry Audio 50+ hours of real infant cry recordings for training infant cry detection, cry classification, and sound event detection models. Manually verified files captured in natural domestic conditions, with per-file metadata on location, background noise, and recording device Contact us and share your feedback — receive additional samples for free! 😊 Key Highlights 50+ hours of real-world… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/infant-cry-detection-dataset.audioaudio-classificationn<1K0 likes85 downloads4mo agoHugging Face24jpdiazpardo /scream_detection_heavy_metal Dataset card for Scream Detection in Heavy Metal Music This dataset contains the processed dataset used in the paper "Scream Detection in Heavy Metal Music" (Kalbag & Lerch, 2022) from the Georgia Institute of Technology. This dataset contains annotations of 57 songs, distributed over 34 bands and 47 albums. The vocal events are labelled into 5 classes: Clean (or sung vocal) Low Fry Scream Mid Fry Scream High Fry Scream Layered Vocals The label "Layered Vocals" has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/jpdiazpardo/scream_detection_heavy_metal.audioaudio-classification1K<n<10K1 likes64 downloads3y agoHugging Face25aegean-ai /engine-anomaly-detection-dataset license: other license_name: ntt license_link: https://zenodo.org/records/3351307/files/LICENSE.pdf?download=1 audio1K<n<10K7 likes61 downloads3y agoHugging Face26rmarcosg /bark-detection Bark detection dataset Dataset Description This dataset comprises both positive and negative samples of audio of 1 second in WAV format, recorded at 44.1kHz. Negative samples include music, voice, claps, whistles and vacuum cleaner noise, among other sound you may record inside a house. Caveats: This is an imbalanced dataset: ~10k negatives vs ~500 positives. Positive samples may include human generated barks. Some (few) positive samples are false positives.… See the full description on the dataset page: https://huggingface.co/datasets/rmarcosg/bark-detection.audioaudio-classification10K<n<100K3 likes58 downloads3y agoHugging Face27MUGEN-Benchmark /Key_Detectionaudion<1K0 likes57 downloads8mo agoHugging Face28speaches-ai /realtime-turn-detection-test-data Realtime speech test recordings Synthetic speech recordings for black-box Realtime API behavior tests in Speaches. Each WAV file is the unmodified output of OpenAI text-to-speech. Tests are responsible for adding silence, combining recordings, and choosing streaming chunk boundaries for their scenarios. metadata.jsonl follows the Hugging Face AudioFolder layout. Each record contains the generation inputs, file digest, expected text, transcription, and word/speech intervals from… See the full description on the dataset page: https://huggingface.co/datasets/speaches-ai/realtime-turn-detection-test-data.audion<1K0 likes54 downloads2mo agoHugging Face29Dhruv0100 /deepfake-audio-detection Deepfake Audio Detection Dataset (v4) Dataset Description This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio. What's New in v4 52% larger: Increased from 1,224 to 1,866 samples (642 new samples) Expanded TTS coverage: Added Hume AI as 6th synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Dhruv0100/deepfake-audio-detection.audioaudio-classification1K<n<10K0 likes53 downloads4d agoHugging Face30PuristanLabs1 /urdu-turn-detection-audio-v2 🗣️ Urdu Turn Detection (Audio Dataset V2) This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech. It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection. 🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.audioaudio-classification10K<n<100K0 likes50 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.