Team Ai
Datasetpublic

batgre/conversational-dynamics-egocom

Conversational Dynamics — EgoCom Derived temporal annotations and model-ready training anchors for conversational-dynamics and turn-taking research, generated from EgoCom. This dataset is produced by the conversational-dynamics-data pipeline. The underlying objective is to expose conversational data in a representation suitable for temporal and action-conditioned models: state_t + action_t → future conversational state Contents Three configurations are provided.… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom.

sourceHugging Faceupdated 2h agoView on Hugging Face
1likes173downloads
Dataset Card

Conversational Dynamics — EgoCom

Derived temporal annotations and model-ready training anchors for conversational-dynamics and turn-taking research, generated from EgoCom.

This dataset is produced by the conversational-dynamics-data pipeline. The underlying objective is to expose conversational data in a representation suitable for temporal and action-conditioned models:

text
state_t + action_t → future conversational state

Contents

Three configurations are provided.

model_ready

One row per candidate temporal anchor. It contains the metadata required to reconstruct past context and future prediction windows without materializing millions of overlapping sequences.

Important fields include:

  • —recording_id
  • —conversation_id
  • —split
  • —anchor_idx
  • —anchor_time
  • —max_context_steps
  • —future_steps
  • —context_valid_ratio
  • —future_valid_ratio
  • —future_event_count
  • —sample_class
  • —is_trainable
  • —schema versions

action_grid

The underlying regular temporal representation used by the anchors.

The current prototype operates at:

text
10 Hz
100 ms per timestep

with focal vocal states:

text
SPEAKING
SILENT
UNKNOWN

and vocal actions:

text
NO_EVENT
ONSET
OFFSET

Invalid or semantically unsupported transitions are masked rather than replaced with artificial labels.

media_manifest

One row per recording, mapping the canonical recording_id to its raw media without shipping the media:

columnmeaning
dataset, recording_idunique key; recording_id equals the action grid's
video_pathvideo file, relative to the corpus root
audio_pathseparate audio file, relative to the corpus root, or null
media_offset_smedia_time_s = decision_time_s + media_offset_s
video_has_audiothe video container carries an audio stream

Paths are relative to the corpus root: the directory holding EgoCom's 240p/ folder (e.g. 240p/20min/<recording_id>.MP4). Join them with your own local copy of the corpus; no absolute path is stored. audio_path is null when the corpus ships no separate audio file: decode the audio track of the video (video_has_audio is true for every recording in this release). For EgoCom the recording clock is the video's own, so media_offset_s is 0.

labels (optional sidecars)

Optional, versioned label sidecars on the same grid, for probing, analysis and experiments. They are plain files, not a datasets config: nothing is read unless it is requested.

text
data/labels/
  registry.json              every label: meaning, modalities, source kind,
                             time reference, validity, storage (164 labels)
  speech/  social/  text/    one directory per extractor:
    manifest.json            versions, configuration digest, input checksums,
                             materialized and unavailable labels
    grid.parquet             the action grid's rows, same order (anchor_row indexes it)
    events.parquet           native-time events (onsets, offsets, floor changes, tokens)
    segments.parquet         native-time intervals (runs, turns, overlaps, silences, ...)
    participants.parquet     per-participant identity and interaction profiles
    recordings.parquet       per-recording participants and metadata

Families: speaker activity per participant, joint state and floor, onsets and offsets with their context, overlaps (within / between / simultaneous), timing with explicit censoring, speech runs and turns, floor transfers with FTO, next-speaker and future targets, profiles, transcript cues. Every label records its source_kind (native_annotation or deterministic here; no model output is presented as ground truth) and a null value always means unknown or censored, never false or zero. Interpretations without a validated source (dialogue acts, addressee, backchannels, social states) are registered as unsupported and not materialized.

Request labels by exact name, family.* or all, optionally filtered by modality, and read only their columns:

python
import pyarrow.parquet as pq
from huggingface_hub import hf_hub_download

path = hf_hub_download(REPO_ID, "data/labels/speech/grid.parquet", repo_type="dataset")
grid = pq.read_table(
    path,
    columns=["recording_id", "decision_index", "participant_ids", "speaker_activity",
             "time_to_next_ego_onset", "time_to_next_ego_onset_valid"],
)

registry.json names each label's file and columns; manifest.json says which labels this corpus materializes and why the others are unavailable.

Definitions, validity semantics and dataset notes: `labels.md` and the generated `labels_registry.md`.

EgoCom's word timings almost never overlap across speakers (27 overlapping word pairs in 157,603 timed words), so overlap labels are nearly empty for this corpus: a property of the transcription, kept as is.

Loading

Model-ready anchors:

python
from datasets import load_dataset

anchors = load_dataset(
    "batgre/conversational-dynamics-egocom",
    "model_ready",
)

Temporal action grid:

python
grid = load_dataset(
    "batgre/conversational-dynamics-egocom",
    "action_grid",
)

model_ready exposes native train, validation and test splits:

python
train = load_dataset(
    "batgre/conversational-dynamics-egocom", "model_ready", split="train"
)
validation = load_dataset(
    "batgre/conversational-dynamics-egocom", "model_ready", split="validation"
)
test = load_dataset(
    "batgre/conversational-dynamics-egocom", "model_ready", split="test"
)

action_grid has a single train split: it is not a set of supervised examples but the complete canonical trajectory the anchors point into.

The split column is kept inside each partition. It is redundant with the file the row lives in, and that is the point: it makes the partitioning a checkable invariant rather than an implicit convention.

Anchors versus trainable anchors

Every partition contains all candidate anchors of its split, including those whose validity ratios fall below the pipeline's thresholds. Use is_trainable to keep only the ones the pipeline considers usable:

splitanchorsof which `is_trainable`
train1 088 7611 072 231
validation86 26885 003
test209 436205 722

The counts in metadata.json refer to the is_trainable subset.

Temporal semantics

An anchor at timestep t represents a possible prediction point. A downstream model may reconstruct:

text
context = [t - L + 1, ..., t]
future  = [t + 1, ..., t + H]

The anchor belongs to the context; prediction starts at t + 1.

For the current release:

text
minimum context: 1 s
maximum context: 5 s
future horizon:  1 s
grid frequency:  10 Hz

Context length is intentionally not fixed by the dataset.

To reconstruct a window, keep the action_grid rows whose recording_id matches the anchor, order them by decision_index, and take the rows around decision_index == anchor_idx.

Splits

Splits are assigned at the conversation level rather than at the anchor level. This prevents synchronized or otherwise related recordings from the same conversation from leaking across training and evaluation splits.

Original dataset splits are preserved when available; split_source records whether an assignment came from the original release or from the pipeline's deterministic seeded fallback.

Provenance and reproducibility

The accompanying metadata.json records the dataset-generation contract and provenance information, including relevant schema versions, temporal geometry, validity thresholds, counts, source checksums and pipeline lineage.

The generating code is maintained separately in the conversational-dynamics-data GitHub repository.

Source data

This repository contains derived temporal annotations only. It does not redistribute EgoCom raw videos or audio; media_manifest only references them by corpus-relative path.

Users requiring the original source media should obtain EgoCom separately and comply with its original terms and licence.

Scope

This release currently represents the vocal state/action layer. The temporal backbone is intended to support additional aligned conversational information in future releases, including audio, visual, gaze, pose, addressee and other multimodal signals.

Model architectures, training loops, sampling strategies and evaluation code are intentionally maintained outside this dataset repository.