datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
real-vs-fake-human-voice-deepfake-audio
Deepfake Audio Dataset
Dataset contains 5,000 audio files, comprising both authentic human recordings and synthetic** AI-generated voice** samples. It designed for advanced research in deepfake detection, focusing on detecting fake voices and generated speech analysis. Specifically engineered to challenge voice authentication systems, it supports the development of robust models for real vs fake human voice recognition.
By utilizing this dataset, researchers and developers can… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/real-vs-fake-human-voice-deepfake-audio.humans-benchmark
HUMANS Benchmark Dataset
Authors: Woody Haosheng Gan¹, William Held²'³, Diyi Yang²
¹University of Southern California, ²Stanford University, ³OpenAthena
This dataset is part of the Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment paper.
HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark is designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/humans-benchmark.Eng-Filipino-Accented-audio-with-human-transcription-call-center-topicThis dataset contains 103+ hours of spontaneous English conversations spoken in a Filipino accent, recorded in a studio environment to ensure crystal-clear audio quality. The conversations are designed as role-play scenarios between agents and customers across a variety of call center domains.
🗣️ Speech Style: Natural, unscripted role-playing between native Filipino-accented English speakers, simulating real-world customer interactions.
🎧 Audio Format: High-quality stereo WAV files, recorded… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Eng-Filipino-Accented-audio-with-human-transcription-call-center-topic.human-robot-conversation-russian
Human-Robot Dataset
The dataset comprises 660+ hours of Russian speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-russian.human-robot-conversation-russian
Human-Robot Conversation Dataset (Russian) - 660+ Hours
Dataset (Russian) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-russian.humans-benchmark
HUMANS Benchmark Dataset (Anonymous, Under Review)
This dataset is part of the HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark, designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned regression weights.
Installation
Install the HUMANS evaluation package from GitHub (our anonymous repo):
# Option 1: Install via pip
pip install… See the full description on the dataset page: https://huggingface.co/datasets/HUMANSBenchmark/humans-benchmark.human-robot-conversation-korean
Human-Robot Conversation Dataset (Korean) - 660+ Hours
Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.az-asr-calls-human-140h
Azerbaijani call-centre speech, human-transcribed
Superseded by cillegio/az-asr-calls-human-496h (2026-10-09).
It contains every clip here unchanged, with identical dev and test splits, plus 356 h more.
Use it for new training. This dataset is kept as-is to reproduce Chinar-F8 v3 and earlier experiments.
Real Azerbaijani call-centre audio with transcripts written by people
listening to it. 29,982 clips, 139.75 h, 8 kHz telephony.
This is the most valuable corpus in the… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-calls-human-140h.az-asr-human-430h
Azerbaijani read speech (LocalDoc + FLEURS)
Azerbaijani read speech assembled from LocalDoc/azerbaijani_asr and
LocalDoc/fleurs-azerbaijani-asr. 371,515 clips, 429.85 h, 16 kHz.
Clean, accurately labelled, and the wrong register for telephony. It is
literature and schoolbooks read aloud, plus FLEURS sentences. Useful for
vocabulary and general acoustics; it will not teach a model what a spontaneous
phone conversation sounds like.
A related experiment is worth knowing about… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-human-430h.human-robot-conversation-korean
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.Thai-H2H-Call-center-audio-with-human-transcriptionThis dataset contains natural Thai-language conversations between human agents and human customers, designed to reflect realistic call center interactions across multiple domains. All conversations are conducted through unscripted role-playing, allowing for spontaneous and dynamic exchanges that closely mirror real-world scenarios.
🗣️ Speech Type: Human-to-human dialogues simulating customer-agent interactions.
🎭 Style: Non-scripted, spontaneous role-playing to capture authentic speech… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Thai-H2H-Call-center-audio-with-human-transcription.real-vs-fake-human-voice-deepfake-audio
Deepfake Audio Dataset - 5,000 Audio
Dataset contains 5,000 audio files, comprising both authentic human recordings and synthetic** AI-generated voice** samples. It designed for advanced research in deepfake detection, focusing on detecting fake voices and generated speech analysis. Specifically engineered to challenge voice authentication systems, it supports the development of robust models for real vs fake human voice recognition. - Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/real-vs-fake-human-voice-deepfake-audio.human-robot-conversation-english
Human-Robot Conversation Dataset (English) - 660+ Hours
Dataset (English) contains 660+ hours of audio featuring dialogues between AI and a human in English across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-english.human-robot-conversation-german
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.human-robot-conversation-english
Human-Robot Dataset
The dataset comprises 660+ hours of English speech across 20,000+ audio files featuring human-robot interactions between AI and humans. It is designed for research in conversational agents, focusing on various speech recognition methods, primarily aimed at advancing language models and machine learning applications.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-english.human-robot-conversation-german
Human-Robot Conversation Dataset (German) - 660+ Hours
Dataset (German) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-german.human-feedback-mentor-sessions-india
license: other
task_categories:
audio-classification
automatic-speech-recognition
text-generation
language:
en
hi
bn
ta
te
ml
mr
or
as
pa
pretty_name: Human Feedback Mentor Sessions India
size_categories:
1K<n<10K
Human Feedback Mentor Sessions India
Dataset Description
A rare and high-value collection of recorded aspirant-mentor interactions
capturing real guidance, feedback, and reasoning corrections in the context
of Indian government exam preparation. Produced by… See the full description on the dataset page: https://huggingface.co/datasets/DataOrigin/human-feedback-mentor-sessions-india.Thai-human-to-machine-call-center-audio-with-scriptThis dataset features natural Thai-language conversations between human speakers and machine agents, simulating real-world call center interactions across a variety of customer service domains. All dialogues are non-scripted and performed as role-play scenarios, capturing spontaneous, realistic exchanges.
🗣️ Speech Type: Human-to-machine conversations, simulating AI agents and human customers' dialogues.
🎭 Style: Spontaneous, unscripted role-playing, designed to reflect actual customer… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Thai-human-to-machine-call-center-audio-with-script.
