Krisp-AI/VoiceIsolation-Benchmark-Dataset
Voice Isolation Benchmark Dataset 265 real-world recordings for measuring how a second voice breaks speech-to-text, and how much Krisp Voice Isolation fixes it. Three scenarios, 47 speakers, real rooms, real headsets. No synthetic mixing. 265 recordings · 47 speakers · 3 scenarios · 65 scripts Why this dataset exists Modern STT engines handle noise well. They still fail when a second person talks near the microphone: they transcribe the wrong speaker, and voice… See the full description on the dataset page: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset.
Voice Isolation Benchmark Dataset
265 real-world recordings for measuring how a second voice breaks speech-to-text, and how much Krisp Voice Isolation fixes it. Three scenarios, 47 speakers, real rooms, real headsets. No synthetic mixing.
<h4><span style="color:#7a54cf">265 recordings · 47 speakers · 3 scenarios · 65 scripts</span></h4>
Why this dataset exists
Modern STT engines handle noise well. They still fail when a second person talks near the microphone: they transcribe the wrong speaker, and voice agents take the turn at the wrong moment. We did not find a public benchmark for this, so we recorded one.
The dataset supports two measurements:
- STT accuracy — Word error rate on raw audio versus the same audio after Voice Isolation, across environments, devices, and noise conditions.
- Perceptual quality — Reference-free speech quality (DNSMOS) before and after processing, on the segments where the primary speaker is talking.
This is not a leaderboard. We are not ranking STT engines. The dataset exists to show the impact of voice isolation under the hardest condition there is: competing voices.
How recordings are structured
Most Work and Call Center recordings follow the same four-step script: the primary speaker begins, a secondary speaker joins, the primary goes silent while the secondary continues alone, then the primary rejoins.
<p align="center"> <svg viewBox="0 0 830 150" width="800" role="img" aria-label="Timeline of one recording: noise, primary, mix, secondary, mix, noise"> <text x="0" y="32" font-family="sans-serif" font-size="13" font-weight="600" fill="#4a4a63">Primary</text> <text x="0" y="78" font-family="sans-serif" font-size="13" font-weight="600" fill="#4a4a63">Secondary</text> <line x1="90" y1="30" x2="800" y2="30" stroke="#dcd6f2" stroke-width="1"/> <line x1="90" y1="76" x2="800" y2="76" stroke="#dcd6f2" stroke-width="1"/> <rect x="770" y="25" width="50" height="10" rx="3" fill="#9aa3b5"/> <rect x="770" y="71" width="50" height="10" rx="3" fill="#9aa3b5"/> <rect x="90" y="25" width="190" height="10" rx="3" fill="#6c47c7"/> <rect x="280" y="25" width="200" height="10" rx="3" fill="#6c47c7"/> <rect x="640" y="25" width="130" height="10" rx="3" fill="#6c47c7"/> <rect x="280" y="71" width="200" height="10" rx="3" fill="#0d9488"/> <rect x="480" y="71" width="290" height="10" rx="3" fill="#0d9488"/> <g font-family="monospace" font-size="11" text-anchor="middle"> <rect x="90" y="112" width="190" height="22" rx="6" fill="#6c47c7" opacity="0.14"/> <text x="185" y="127" fill="#141428">primary</text> <rect x="280" y="112" width="200" height="22" rx="6" fill="#e0932b" opacity="0.18"/> <text x="380" y="127" fill="#141428">mix</text> <rect x="480" y="112" width="160" height="22" rx="6" fill="#0d9488" opacity="0.16"/> <text x="560" y="127" fill="#141428">secondary</text> <rect x="640" y="112" width="130" height="22" rx="6" fill="#e0932b" opacity="0.18"/> <text x="705" y="127" fill="#141428">mix</text> <rect x="770" y="112" width="50" height="22" rx="6" fill="#9aa3b5" opacity="0.18"/> <text x="795" y="127" fill="#141428">noise</text> </g> <g stroke="#dcd6f2" stroke-dasharray="3 3"> <line x1="280" y1="12" x2="280" y2="108"/><line x1="480" y1="12" x2="480" y2="108"/><line x1="640" y1="12" x2="640" y2="108"/><line x1="770" y1="12" x2="770" y2="108"/> </g> <text x="820" y="100" text-anchor="end" font-family="sans-serif" font-size="11" fill="#8a8aa3">time →</text> </svg> </p>
<sub>One recording, four labeled segments. The "secondary" window is the stress test: the correct transcript for that stretch is empty, and most STT engines transcribe the background voice anyway.</sub>
This forces every system through speaker transitions, overlapping speech, and the stretch where only the background voice is present.
Phone Calls do not follow this structure. They contain only the primary speaker. There is no deliberate second voice; the background is ambient babble and environmental noise from the car, street, or shop. That set tests something else: whether voice isolation stays safe on distorted, band-limited audio where there is little to remove.
The three scenarios
All recordings were made by the Krisp team in real environments, on the hardware people actually use. For the call-center set we recorded inside active call centers during their normal workflow. Nothing was mixed or augmented afterwards.
Work
<h4><span style="color:#7a54cf">96 recordings · 29 speakers · 24 devices · 4 room types · 17 scripts</span></h4>
Typical work-from-home and office-meeting situations. Recorded in furnished small rooms (under 10 m²), medium rooms (10–20 m²), large rooms (over 20 m²), and open-space offices. The secondary voice is a colleague, a family member, or someone sharing the space. 29 speakers (15 female, 14 male) on 24 audio devices, from professional headsets to consumer earbuds and gaming headsets. Scripts are longer and denser than the call-center set: proper names, numbers, dates, and terms.
Audio devices used<br> Jabra Evolve 10, Jabra Evolve 20 MS, Jabra Evolve2 40, Jabra Evolve 75se, Jabra Engage 50, Poly Blackwire 8225, Poly Voyager Focus, Mpow HC5, Apple AirPods, Apple EarPods, Bose QC35, JBL Tune 500BT, JBL Tune 710BT, Soundcore Life Q30, Soundcore P25i, Samsung Galaxy Buds Live, Redmi Buds 3 Lite, Logitech G435, Logitech G733, Logitech Zone Vibe 100, Razer BlackShark V3 Pro, Livey 805DM, Cyber Acoustics AC-204TR
Call Center
<h4><span style="color:#7a54cf">106 recordings · 7 speakers · 13 headsets · 4 locations · 28 scripts</span></h4>
Recorded inside real call center environments across 4 distinct locations, each contributing its own background noise profile — ambient chatter from neighboring agents, keyboard typing, phone rings, and general office activity. 7 speakers (4 female, 3 male) on professional headsets. 28 scripts cover agent–customer calls in banking, insurance, healthcare, IT support, telecom, e-commerce, travel, and more. Only the agent side is recorded.
Audio devices used<br> Jabra Biz 1500, Jabra Evolve 10, Jabra Evolve2 30, Jabra Evolve2 40, Cyber Acoustics AC-204TR, Cyber Acoustics AC-304TR, Livey 410DM Plus, Livey 500DM AINC, Livey 805DM, Logitech H340, Poly Blackwire 8225, Telekonnectors FS V2, Telekonnectors Sirius QD
Phone Calls
<h4><span style="color:#7a54cf">63 recordings · 26 speakers · 18 phones · 4 location types · 20 scripts</span></h4>
Speakers on the move: driving, walking outdoors, commuting, or inside shops and cafés. Most recordings come from cars, the most common mobile calling environment. The challenge is the signal itself, not a second speaker: Bluetooth car microphones, phones on speaker far from the mouth, earbuds in a moving vehicle, telephony codecs, and low bandwidth on top of street noise, music, cabin rumble, and crowd chatter. 26 speakers (10 female, 16 male) on 18 phone models, recorded from 17 car models as well as on foot and in public spaces.
Phones used<br> iPhone 13, iPhone 13 Pro Max, iPhone 14, iPhone 14 Pro, iPhone 15, iPhone 15 Pro, iPhone 15 Pro Max, iPhone 16, iPhone 16 Pro, iPhone 17, iPhone 17 Pro, Samsung Galaxy S7, Samsung Galaxy S25, Samsung Fold 7, Sony Xperia 5 IV, Redmi 13C, Honor 400 Pro, OnePlus 12
Annotations
Every recording ships with a hand-written transcript, verified by human annotators. Transcripts capture speech exactly as spoken: hesitations, false starts, and self-corrections stay in. Partial words and restarts such as ano.., astonishing or rec.. resolved are kept, so ground truth is what was said, not a cleaned-up version.
Audio segment labels
- <span style="color:#6c47c7; font-size:0.7em">●</span> primary — target speaker (may include ambient bubble noise)
- <span style="color:#e0932b; font-size:0.7em">●</span> mix — both speakers overlapping
- <span style="color:#0d9488; font-size:0.7em">●</span> secondary — secondary speaker only
- <span style="color:#9aa3b5; font-size:0.7em">●</span> noise — no speech, only background noise
Segment labels let you measure STT and voice isolation under each condition on its own, not just on the recording as a whole.
Audio Format
All recordings are uncompressed WAV files (PCM 16-bit signed little-endian), mono channel, sampled at 32 kHz.
Metadata fields
Each scenario has a metadata.jsonl file with one entry per recording. All scenarios share these fields:
file_name— path to the audio fileid— unique sample identifierspeaker_id— anonymized speaker labelgender—female/malefull_transcript— complete ground-truth transcriptsegments— list of time-stamped objects, each with start and end time, transcript text, and the segment labelrecording_device— headset (Call Center), audio device (Work), or microphone type (Phone Calls)environment— call-center location, room type, or calling context
Phone Calls add two fields:
phone— the phone modelcar_model— the car model, when recorded from a car
What is released
The full dataset is published on Hugging Face as a Collection: the recordings, transcripts, and metadata, plus the output of each Voice Isolation model on every file. Benchmark results across 11 STT configurations are published as a Space, with the methodology (corpus-level WER via jiwer, DNSMOS-C NISQA on primary segments) documented there.
<sub>Krisp Technologies, Inc. · Recordings are anonymized; speaker IDs carry no personal information.</sub>
