batgre/conversational-dynamics-egocom
Conversational Dynamics — EgoCom Derived temporal annotations and model-ready training anchors for conversational-dynamics and turn-taking research, generated from EgoCom. This dataset is produced by the conversational-dynamics-data pipeline. The underlying objective is to expose conversational data in a representation suitable for temporal and action-conditioned models: state_t + action_t → future conversational state Contents Three configurations are provided.… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom.
Conversational Dynamics — EgoCom
Derived temporal annotations and model-ready training anchors for conversational-dynamics and turn-taking research, generated from EgoCom.
This dataset is produced by the conversational-dynamics-data pipeline. The underlying objective is to expose conversational data in a representation suitable for temporal and action-conditioned models:
state_t + action_t → future conversational stateContents
Three configurations are provided.
model_ready
One row per candidate temporal anchor. It contains the metadata required to reconstruct past context and future prediction windows without materializing millions of overlapping sequences.
Important fields include:
recording_idconversation_idsplitanchor_idxanchor_timemax_context_stepsfuture_stepscontext_valid_ratiofuture_valid_ratiofuture_event_countsample_classis_trainable- schema versions
action_grid
The underlying regular temporal representation used by the anchors.
The current prototype operates at:
10 Hz
100 ms per timestepwith focal vocal states:
SPEAKING
SILENT
UNKNOWNand vocal actions:
NO_EVENT
ONSET
OFFSETInvalid or semantically unsupported transitions are masked rather than replaced with artificial labels.
media_manifest
One row per recording, mapping the canonical recording_id to its raw media without shipping the media:
Paths are relative to the corpus root: the directory holding EgoCom's 240p/ folder (e.g. 240p/20min/<recording_id>.MP4). Join them with your own local copy of the corpus; no absolute path is stored. audio_path is null when the corpus ships no separate audio file: decode the audio track of the video (video_has_audio is true for every recording in this release). For EgoCom the recording clock is the video's own, so media_offset_s is 0.
labels (optional sidecars)
Optional, versioned label sidecars on the same grid, for probing, analysis and experiments. They are plain files, not a datasets config: nothing is read unless it is requested.
data/labels/
registry.json every label: meaning, modalities, source kind,
time reference, validity, storage (164 labels)
speech/ social/ text/ one directory per extractor:
manifest.json versions, configuration digest, input checksums,
materialized and unavailable labels
grid.parquet the action grid's rows, same order (anchor_row indexes it)
events.parquet native-time events (onsets, offsets, floor changes, tokens)
segments.parquet native-time intervals (runs, turns, overlaps, silences, ...)
participants.parquet per-participant identity and interaction profiles
recordings.parquet per-recording participants and metadataFamilies: speaker activity per participant, joint state and floor, onsets and offsets with their context, overlaps (within / between / simultaneous), timing with explicit censoring, speech runs and turns, floor transfers with FTO, next-speaker and future targets, profiles, transcript cues. Every label records its source_kind (native_annotation or deterministic here; no model output is presented as ground truth) and a null value always means unknown or censored, never false or zero. Interpretations without a validated source (dialogue acts, addressee, backchannels, social states) are registered as unsupported and not materialized.
Request labels by exact name, family.* or all, optionally filtered by modality, and read only their columns:
import pyarrow.parquet as pq
from huggingface_hub import hf_hub_download
path = hf_hub_download(REPO_ID, "data/labels/speech/grid.parquet", repo_type="dataset")
grid = pq.read_table(
path,
columns=["recording_id", "decision_index", "participant_ids", "speaker_activity",
"time_to_next_ego_onset", "time_to_next_ego_onset_valid"],
)registry.json names each label's file and columns; manifest.json says which labels this corpus materializes and why the others are unavailable.
Definitions, validity semantics and dataset notes: `labels.md` and the generated `labels_registry.md`.
EgoCom's word timings almost never overlap across speakers (27 overlapping word pairs in 157,603 timed words), so overlap labels are nearly empty for this corpus: a property of the transcription, kept as is.
Loading
Model-ready anchors:
from datasets import load_dataset
anchors = load_dataset(
"batgre/conversational-dynamics-egocom",
"model_ready",
)Temporal action grid:
grid = load_dataset(
"batgre/conversational-dynamics-egocom",
"action_grid",
)model_ready exposes native train, validation and test splits:
train = load_dataset(
"batgre/conversational-dynamics-egocom", "model_ready", split="train"
)
validation = load_dataset(
"batgre/conversational-dynamics-egocom", "model_ready", split="validation"
)
test = load_dataset(
"batgre/conversational-dynamics-egocom", "model_ready", split="test"
)action_grid has a single train split: it is not a set of supervised examples but the complete canonical trajectory the anchors point into.
The split column is kept inside each partition. It is redundant with the file the row lives in, and that is the point: it makes the partitioning a checkable invariant rather than an implicit convention.
Anchors versus trainable anchors
Every partition contains all candidate anchors of its split, including those whose validity ratios fall below the pipeline's thresholds. Use is_trainable to keep only the ones the pipeline considers usable:
The counts in metadata.json refer to the is_trainable subset.
Temporal semantics
An anchor at timestep t represents a possible prediction point. A downstream model may reconstruct:
context = [t - L + 1, ..., t]
future = [t + 1, ..., t + H]The anchor belongs to the context; prediction starts at t + 1.
For the current release:
minimum context: 1 s
maximum context: 5 s
future horizon: 1 s
grid frequency: 10 HzContext length is intentionally not fixed by the dataset.
To reconstruct a window, keep the action_grid rows whose recording_id matches the anchor, order them by decision_index, and take the rows around decision_index == anchor_idx.
Splits
Splits are assigned at the conversation level rather than at the anchor level. This prevents synchronized or otherwise related recordings from the same conversation from leaking across training and evaluation splits.
Original dataset splits are preserved when available; split_source records whether an assignment came from the original release or from the pipeline's deterministic seeded fallback.
Provenance and reproducibility
The accompanying metadata.json records the dataset-generation contract and provenance information, including relevant schema versions, temporal geometry, validity thresholds, counts, source checksums and pipeline lineage.
The generating code is maintained separately in the conversational-dynamics-data GitHub repository.
Source data
This repository contains derived temporal annotations only. It does not redistribute EgoCom raw videos or audio; media_manifest only references them by corpus-relative path.
Users requiring the original source media should obtain EgoCom separately and comply with its original terms and licence.
Scope
This release currently represents the vocal state/action layer. The temporal backbone is intended to support additional aligned conversational information in future releases, including audio, visual, gaze, pose, addressee and other multimodal signals.
Model architectures, training loops, sampling strategies and evaluation code are intentionally maintained outside this dataset repository.
