full-duplex
otoSpeech-full-duplex-turn-104h
Dataset Card for otoSpeech-full-duplex-turn-104h
Contact
Website: https://oto.earthEmail: agent@oto.earth
Dataset Summary
otoSpeech-full-duplex-turn-104h is an English, full-duplex conversational speech dataset for research on turn-taking and related spoken-dialogue phenomena. It contains 420 two-speaker conversations totaling approximately 104.94 hours. Each conversation includes time-aligned, channel-separated audio, a stereo combined recording… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-turn-104h.otoSpeech-full-duplex-280h
📢 Notice: We released a new processed version of otoSpeech. Available here: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-processed-141h
Dataset Card for otoSpeech-full-duplex-280h: Full-Duplex Conversational Speech Dataset
Contact
Website: https://oto.earth
Mail: consome@oto.earth
Dataset Summary
otoSpeech-full-duplex-280h is a 280-hour, full-duplex, two-speaker conversational speech dataset.
Each sample includes 48 kHz… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-280h.otoSpeech-full-duplex-processed-141h
Dataset Card for otoSpeech-full-duplex-processed-141h: Full-Duplex Conversational Speech Dataset
Dataset Summary
otoSpeech-full-duplex-processed-141h is a full-duplex, two-speaker conversational speech dataset. It is derived from otoSpeech-full-duplex-280h and has been curated and processed as follows:
Selected high-quality conversations based on human reviews.
Applied noise reduction and speech enhancement to improve audio quality.
Added new samples collected after the… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-processed-141h.video-full-duplex-benchmark
VideoFDB: Video-Full-Duplex-Benchmark
Project Page · HuggingFace · Paper (arXiv)
Dataset Description
A benchmark dataset of annotated, two-person video conference recordings designed to support the evaluation of multimodal AI agents in conversational settings. The dataset covers 11 distinct conversational dynamics — spanning verbal, nonverbal, and mixed-modality behavior — annotated through a three-pass human-in-the-loop pipeline.
The benchmark consists of trimmed… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-full-duplex-benchmark.Full-Duplex-Bench-Data214-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Dataset
Description
주제 기반 대화로 수집한 한국어(대한민국) 멀티스트림 자연 발화 스마트폰 음성 데이터셋입니다. 각 음성 데이터에는 대화 내용을 전사한 텍스트와 화자 ID, 성별, 연령 등의 속성 정보가 포함되어 있습니다. 본 데이터셋은 다양한 지역과 배경을 가진 폭넓은 화자들로부터 수집되었으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1704?source=hf.kr
Specifications
Format
16 kHz, 16 bit, WAV, 모노 채널
Content category
정해진 주제 없이 자유롭게 진행된 자연 대화
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/214-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Dataset.
