Team Ai
Datasetpublic

tugrulbayrak/Real-TurnTurk

Real-TurnTurk English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
3likes515downloads
Dataset Card

Real-TurnTurk

![arXiv](https://arxiv.org/abs/2608.22071)

English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the dataset ships frame-level facial feature vectors (head pose, gaze, 52 ARKit blendshapes), time-aligned transcripts, and frame-level speech activity annotations on a shared 62.5 ms grid.

Türkçe: Real-TurnTurk, sesli diyalog sistemlerinde "sıra alma / sıra değişimi" (turn-taking) tahminini geliştirmek amacıyla oluşturulmuş çok kipli (multimodal) Türkçe iki kanallı konuşma veri kümesidir. Oluşturduğumuz diğer veri kümesi Syn-TurnTurk'ten farklı olarak buradaki tüm diyaloglar, iki kişi arasında görüntülü görüşme üzerinden kaydedilmiş gerçek ve senaryosuz konuşmalardır. Her katılımcı ayrı bir ses kanalında kaydedilmiştir; bu sayede konuşmacı ataması kesindir ve herhangi bir ayrıştırma (diarization) modeli gerektirmez. Ses kayıtlarının yanı sıra veri kümesi kare düzeyinde yüz öznitelik vektörleri (kafa pozu, bakış, 52 ARKit blendshape), zaman hizalı transkriptler ve ortak 62.5 ms ızgarada kare düzeyinde konuşma aktivitesi etiketleri içerir.


Use Cases / Kullanım Alanları

English: This dataset can be used for the following AI tasks:

  • —Turn-taking prediction in Turkish spoken dialogue systems
  • —Multimodal speech analysis
  • —Time-aligned transcription and voice activity detection (VAD)
  • —Speech modeling supported by facial feature vectors (ARKit blendshapes, head pose)

Türkçe: Bu veri kümesi aşağıdaki yapay zeka görevleri için kullanılabilir:

  • —Türkçe diyalog sistemlerinde sıra değişimi ve sıra alma (turn-taking prediction) tahmini
  • —Çok kipli (multimodal) konuşma analizi
  • —Zaman hizalı transkript ve ses aktivite tespiti (VAD)
  • —Yüz öznitelik vektörleri (ARKit blendshape, kafa pozu) ile desteklenen konuşma modellemesi

Dataset Summary / Veri Kümesi Özeti

Metric / MetrikValue / Değer
Conversations / Diyalog6
Unique Speakers / Farklı Konuşmacı6 (p01-p06)
Speaker Channels / Konuşmacı Kanalı12
Total Duration / Toplam Süre1.94 h (6,987 s)
Speaker-Channel Audio / Konuşmacı-Kanal Ses3.88 h
Turns / Konuşma Sırası1,642
Backchannels / Geri Bildirim360
Turn Transitions / Sıra Değişimi800
Transcript Segments / Transkript Satırı2,502
Transcript Words / Transkript Kelime16,027
Facial Feature Rows / Yüz Öznitelik Satırı223,602
Frame Rate / Kare Hızı16 fps (62.5 ms)
Language / DilTurkish / Türkçe

Speakers recur across recordings, so person-level comparison is possible: p01 and p06 each appear in three conversations, p03 and p04 in two. / Konuşmacılar kayıtlar arasında tekrar ettiği için kişi bazlı karşılaştırma mümkündür: p01 ve p06 üçer, p03 ve p04 ikişer diyalogda yer alır.


Files / Dosyalar

Each recording lives in its own directory named by recording ID. / Her kayıt, kayıt kimliğiyle adlandırılmış kendi dizinindedir.

<recording_id>/
├── <recording_id>__L_<person>.webm      # left  panel speaker, isolated audio
├── <recording_id>__R_<person>.webm      # right panel speaker, isolated audio
├── <recording_id>__text.csv             # transcript with speaker + timing
├── <recording_id>__speech_activity.csv  # frame-level voice activity
├── <recording_id>__video_features.csv   # frame-level facial features
└── <recording_id>__animation.mp4        # rendered face-vector animation
File / DosyaContent / İçerik
*__L_*.webm, *__R_*.webmPer-speaker isolated audio. Cross-channel bleed is 73-78 dB below the speaker's own level — the opposite speaker sits under the noise floor. / Konuşmacı başına yalıtılmış ses. Kanallar arası sızma, kişinin kendi seviyesinin 73-78 dB altında; karşı konuşmacı gürültü tabanının altında kalıyor.
*__text.csvTranscript, one row per segment. / Transkript, segment başına bir satır.
*__speech_activity.csvVoice activity per person per frame (Silero VAD on the isolated channels). / Kişi ve kare başına konuşma aktivitesi (yalıtılmış kanallar üzerinde Silero VAD).
*__video_features.csv71 columns per person per frame — head pose, eye/mouth geometry, gaze, 52 ARKit blendshapes. / Kişi ve kare başına 71 kolon — kafa pozu, göz/ağız geometrisi, bakış, 52 ARKit blendshape.
*__animation.mp4Face vectors rendered as animation, for visual inspection. / Yüz vektörlerinin animasyona dönüştürülmüş hâli, görsel denetim için.

Column Schemas / Kolon Şemaları

text.csv

ColumnDescription
start_time, end_timeHH:MM:SS.mmm
speakerp01-p06
textTranscribed utterance / Konuşma metni

speech_activity.csv

ColumnDescription
recordingRecording ID / Kayıt kimliği
frame, t_secFrame index and time, 62.5 ms step / Kare indeksi ve zaman, 62.5 ms adım
personp01-p06
speaking0 / 1

VAD settings: min_silence_duration_ms=300, speech_pad_ms=100, min_speech_duration_ms=120.

video_features.csv

Column groupUnit / Birim
recording, frame, t_sec, t_msidentifiers and timestamps / kimlik ve zaman damgaları
person, panelperson code, source tile (L/R) / kişi kodu, kaynak karo
face_detected0/1 — measurement columns are empty when 0 / 0 ise ölçüm kolonları boş
pitch_deg, yaw_deg, roll_deghead angles, degrees / kafa açıları, derece
eye_open_l, eye_open_reye aperture ratio (vertical/horizontal) / göz açıklığı oranı
mouth_open, mouth_widthlip aperture and mouth width ratios / dudak açıklığı ve ağız genişliği oranları
eye_dist_pxouter eye-corner distance, px — camera distance proxy / dış göz köşeleri arası piksel
gaze_xhorizontal gaze, + = person's right / yatay bakış, + = kişinin sağı
tx, ty, tzhead translation, MediaPipe units (relative, not metric) / kafa ötelemesi, MediaPipe birimi (bağıl)
52 blendshapesARKit set, 0..1

face_detected=0 rows are kept, with measurement columns left empty, so the time axis stays continuous and gaps remain visible. / face_detected=0 satırları silinmemiştir; ölçüm kolonları boş bırakılmıştır, böylece zaman ekseni kesintisiz kalır ve boşluklar görünür olur.


Technical Specifications / Teknik Özellikler

Audio / Ses — *__L_*.webm, *__R_*.webm

Property / ÖzellikValue / Değer
CodecOpus
Sample rate / Örnekleme hızı48 kHz
Channels / Kanal1 (mono)
Bitrate / Bit hızı~126-128 kbps

Speech activity annotations were computed after downmixing to 16 kHz mono. / Konuşma aktivitesi etiketleri 16 kHz mono'ya indirildikten sonra hesaplanmıştır.

Animation / Animasyon — *__animation.mp4

Property / ÖzellikValue / Değer
VideoH.264, yuv420p, 1920×1080
Frame rate / Kare hızı16 fps, constant / sabit
Bitrate / Bit hızı~3.2 Mbps
Audio / SesAAC, 16 kHz mono — mixed conversation audio, as in the source video / kaynak videodaki karışık konuşma sesi

Per-speaker isolated audio is in the .webm files. / Konuşmacı bazlı yalıtılmış ses .webm dosyalarındadır.

Feature Extraction / Öznitelik Çıkarımı

Property / ÖzellikValue / Değer
Source video / Kaynak videoH.264, 1920×1080, 16 fps
Layout / Düzenside-by-side dyad, 960×1080 per participant tile (panel = L/R) / yan yana ikili düzen, katılımcı başına 960×1080 karo
LandmarkerMediaPipe Face Landmarker 0.10.35, face_landmarker.task (float16)
Mode / ModRunningMode.VIDEO, num_faces=1, CPU
Mesh478 points + 52 ARKit blendshapes / 478 nokta + 52 ARKit blendshape
Sampling / Örnekleme1:1 with source frames, no interpolation → 62.5 ms step / kaynak kareyle 1:1, interpolasyon yok → 62.5 ms adım

Landmark coordinates are normalised to 0..1 relative to the participant tile, not pixels; video_features.csv holds measurements derived from them. The source video.mp4 and the raw 478-point landmarks are not part of this release. / Landmark koordinatları piksel değil, katılımcı karosuna göre 0..1 normalize edilmiştir; video_features.csv bunlardan türetilmiş ölçümleri taşır. Kaynak video.mp4 ve ham 478 nokta landmark bu yayına dahil değildir.

Size / Boyut

Total 3.0 GB, 400-693 MB per recording; animation.mp4 accounts for roughly 80% of the volume. / Toplam 3.0 GB, kayıt başına 400-693 MB; hacmin yaklaşık %80'i animation.mp4 dosyalarından geliyor.


Statistical Analysis / İstatistiksel Analizler

1. Floor Transfer Offset (FTO)

FTO is the gap between the end of one speaker's turn and the start of the next speaker's turn. Negative values mean the next speaker started before the previous one finished. Backchannels are excluded. / FTO, bir konuşmacının sırasının bitişi ile diğerinin başlangıcı arasındaki farktır. Negatif değer, sonraki konuşmacının öncekinin sözü bitmeden başladığını gösterir. Geri bildirimler hesaba katılmamıştır.

Metric / MetrikAll / Tümü\FTO\≤ 3 s (94%)
Mean / Ortalama−0.116 s+0.081 s
Median / Medyan+0.188 s+0.188 s
Std Dev / Standart Sapma2.599 s0.927 s
P10—−1.250 s
P25—−0.438 s
P75—+0.625 s
P90—+1.125 s
Negative (overlapping transitions) / Negatif (örtüşmeli devir)42.1%—

2. Turn Durations / Sıra Süreleri

Metric / MetrikValue / Değer
Mean / Ortalama3.955 s
Median / Medyan2.438 s
Std Dev / Standart Sapma4.352 s
Max / En Uzun35.9 s

3. Per-Recording Breakdown / Kayıt Bazlı Dağılım

RecordingSpeakersDurationTurnsBackchannelsTransitionsMedian FTOMedian TurnOverlapSilence
rec01p06 (L), p03 (R)959.8 s2153373+0.250 s2.69 s67 s122 s
rec02p04 (L), p06 (R)1027.9 s2282678+0.438 s2.75 s60 s159 s
rec03p03 (L), p01 (R)1464.4 s37799228+0.093 s2.12 s350 s149 s
rec04p01 (L), p02 (R)982.5 s31198168−0.062 s2.06 s255 s129 s
rec05p05 (L), p06 (R)1529.8 s23549118+0.188 s4.81 s105 s124 s
rec06p04 (L), p01 (R)1023.1 s27655135+0.188 s2.28 s152 s119 s

Conversational style varies widely across pairs: rec02 is orderly (median FTO +438 ms, 3% simultaneous speech) while rec04 and rec03 are lively and overlap-heavy (12-13% simultaneous speech). / Çiftler arasında konuşma tarzı belirgin şekilde değişiyor: rec02 sıralı ilerliyor (medyan FTO +438 ms, %3 eşzamanlı konuşma), rec04 ve rec03 ise canlı ve örtüşme yoğun (%12-13 eşzamanlı konuşma).

4. Overlap and Silence / Örtüşme ve Sessizlik

Metric / MetrikValue / Değer
Overlapping blocks / Örtüşen blok1,200
Total overlap / Toplam örtüşme989.3 s
Simultaneous speech / Eşzamanlı konuşma491.4 s (7.0%)
Mutual silence / Karşılıklı sessizlik801.4 s (11.5%)

5. Speech Time per Speaker / Konuşmacı Başına Konuşma Süresi

SpeakerRecordingsSpeech / Konuşma
p0132,035 s
p0632,084 s
p032868 s
p042643 s
p021509 s
p051539 s

Measured Audio-Visual Lead / Ölçülen Ses-Görüntü Öncelemesi

Mouth movement precedes voicing: speakers open and position the articulators before phonation begins. This lead was measured per channel by cross-correlating mouth-movement energy (|Δ mouth_open|, 7-frame smoothing) against the isolated-channel VAD. / Ağız hareketi seslendirmeden önce başlar: konuşmacı, ses çıkmadan önce artikülatörleri konumlandırır. Bu önceleme, ağız hareket enerjisi (|Δ mouth_open|, 7 kare yumuşatma) ile yalıtılmış kanal VAD'i çapraz korelasyona sokularak kanal başına ölçülmüştür.

RecordingPersonLead / Öncelemer
rec01p06+438 ms0.677
rec01p03+438 ms0.588
rec02p04+125 ms0.522
rec02p06+438 ms0.581
rec03p03+438 ms0.654
rec03p01+375 ms0.566
rec04p01+375 ms0.619
rec04p02+375 ms0.496
rec05p05+1125 ms0.693
rec05p06+688 ms0.669
rec06p04+188 ms0.624
rec06p01+312 ms0.576

The lead is partly person-specific and reproducible across recordings: p04 measures +125 and +188 ms in two different conversations, p03 +438 ms in both of its recordings, p01 +375/+375/+312 ms across three. rec05 sits highest on both channels. / Önceleme kısmen kişiye özgü ve kayıtlar arasında tekrarlanabilir: p04 iki ayrı diyalogda +125 ve +188 ms, p03 iki kaydında da +438 ms, p01 üç kaydında +375/+375/+312 ms ölçülüyor. rec05 her iki kanalında da en yüksek değerlere sahip.

Anyone fusing face and audio at sub-second resolution should account for these per-channel values, since the lead is comparable in magnitude to the FTO distribution being modelled. / Yüz ve sesi saniye altı çözünürlükte birleştirecek olanlar bu kanal bazlı değerleri hesaba katmalıdır; önceleme, modellenen FTO dağılımıyla aynı büyüklük mertebesindedir.


Contributors / Katkıda Bulunanlar

License / Lisans

This dataset is licensed under the Apache 2.0 License. / Bu veri seti Apache 2.0 Lisansı ile sunulmaktadır.


Citation / Atıf

If you use this dataset, please cite the accompanying paper. / Bu veri kümesini kullanıyorsanız lütfen ilgili bildiriye atıf veriniz.

bibtex
@misc{bayrak2026realturnturkmultimodalturkishcorpus,
      title={Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction}, 
      author={Ahmet Tuğrul Bayrak and Fatma Nur Korkmaz and Bekir Berker Türker and Mustafa Sertaç Türkel and Alper Kaplan},
      year={2026},
      eprint={2608.22071},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2608.22071}, 
}