JSALT2026-Conv-AI-Simulator/turnbench-dev-no-backchannel
TurnBench Dev - Backchannels Removed A derivative of mundo-ai/turn-benchmark-dev with every majority-annotated backchannel removed from the audio: 1853 backchannels across 38 conversations, 2077.0 seconds in total, cut out of the speaker's own channel and replaced by background noise taken from elsewhere in that same channel. Everything else is the original recording, sample for sample. Same conversations, same duration, same timeline, same annotator tracks, same speech -- only… See the full description on the dataset page: https://huggingface.co/datasets/JSALT2026-Conv-AI-Simulator/turnbench-dev-no-backchannel.
TurnBench Dev - Backchannels Removed
A derivative of `mundo-ai/turn-benchmark-dev` with every majority-annotated backchannel removed from the audio: 1853 backchannels across 38 conversations, 2077.0 seconds in total, cut out of the speaker's own channel and replaced by background noise taken from elsewhere in that same channel.
Everything else is the original recording, sample for sample. Same conversations, same duration, same timeline, same annotator tracks, same speech -- only the backchannels are gone. That makes the pair (original, this) a controlled A/B: a difference in a model's behaviour between the two comes from the backchannels and from nothing else.
Why replace rather than mute
Cutting a backchannel out shortens the recording and shifts every timestamp after it; muting one leaves a hole of digital silence, which no microphone has ever recorded and which a model that listens to a channel's silence structure can spot immediately. Filling the hole with the same channel's own room tone, at the level of the background around it, leaves an intact conversation in which the listener simply did not say "mm-hmm".
What counts as a backchannel
The three annotator tracks the source ships (speaker_{1,2}_annotation_{a,b,c}), voted per moment of audio: a moment is a backchannel where at least 2 of 3 annotators labelled it Continuer Backchannel, Acknowledgement Backchannel, Reaction Backchannel. Bounded Response -- a short but floor-taking reply -- is not a backchannel here, which follows the source's own label set.
That rule reproduces the is_backchannel field of the dataset's shipped merged annotation dump exactly: all 823 of its backchannel segments are covered by this vote and none of its 2420 non-backchannel segments is. Voting on the raw tracks, rather than reading the dump, is what extends the removal to all 38 dev conversations instead of the 12 the dump covers.
Adjacent voted spans within 0.15s are removed as one, spans shorter than 0.05s are ignored, and each removal is padded by 0.03s either side so the 0.01s crossfade at its edges lands on the pad rather than on the first or last syllable.
Where the noise comes from
Per channel, from that channel only -- the background floors here span roughly -38 to -70 dBFS across channels, so noise borrowed from another channel would be an audible level step. Candidates are taken in tiers, first tier with enough material:
Every candidate stretch is eroded by 0.25s at both ends, because a gap between two turns is mostly the decay of one and the breath before the next. Channel Bleed -- the partner audible in this channel -- is never used as noise.
Which chunk fills which hole is decided per hole, and on this material that matters more than the tier does: one channel's background drifts by up to 20 dB over twelve minutes, so no single channel-wide "noise floor" is right for the whole recording. For each hole, the target level is the median of the background nearest to it in time (2.0s of material); candidates are the chunks within 6.0 dB of that target, nearest in time first (this is what keeps an unannotated cough or a chair scrape out of the fill); and the drawn filler is gain-matched to the target, capped at +-15.0 dB. Holes longer than the available chunks are covered by concatenating several with crossfades. Each entry in speaker_{1,2}_backchannel_removals carries level_vs_local_background_db, the residual after all of that -- what is left of the level difference between the fill and its surroundings.
Columns
Same schema as the source, with two changes.
- Added:
speaker_{1,2}_backchannel_removals, one entry per removal, withstart_s,end_s,votes,labels, the removedtext, the level before and after, the local background level, the gain applied to the filler, and thenoise_sources(start/end in the same channel) it was built from. This is the complete record of what was taken out and what replaced it. - Dropped:
speaker_{1,2}_audio_preview. Those are Opus previews of the original audio and would still contain every backchannel.
speaker_{1,2}_audio stays 48 kHz mono 24-bit FLAC, and samples outside the removals are bit-identical to the source. conversation_id, metadata and audio_status pass through untouched.
The annotator tracks are passed through unchanged and still describe the original audio. Their backchannel labels are now the record of what was removed: they mark moments where the speaker no longer says anything. Everything else in them -- turns, interruptions, transcripts -- still lines up with the audio.
Quickstart
from datasets import load_dataset
ds = load_dataset("<this-repo>", split="dev")
row = ds[0]
print(row["conversation_id"], len(row["speaker_1_backchannel_removals"]), "removals")
for removal in row["speaker_1_backchannel_removals"][:3]:
print(removal["start_s"], removal["end_s"], removal["votes"], removal["text"])Per-conversation removals
Known limits
- Cross-channel bleed is not touched. Where the two speakers were recorded in the same room, a backchannel removed from its own channel can still be faintly present in the partner's channel as bleed. Measured over the removed spans in which the partner is themselves unannotated, that residue sits at -63.0 dBFS RMS against a partner-channel floor of -68.1 dBFS (2.5 dB above it) in the median conversation. Attenuating the partner's channel was rejected: it would risk cutting that speaker's real words, and the removal is defined per channel.
backchannel_removals.jsonreports the figure for every conversation undercrosstalk. - The annotation is the ground truth, including its misses. A backchannel that fewer than 2 annotators caught is still in the audio.
- The `audio_status: noisy` conversations are the hard ones. They are also the ones annotated end to end with no silent stretch left over, so they are where the noise source falls back to a lower tier. Their fillers are the ones worth listening to first.
- Verified per channel by
verify_turnbench_no_backchannel: identical timeline, bit-exact outside the removals, every removal at the local background level, no crossfade clicks, complete coverage of the voted spans, and no non-backchannel speech removed.
