naplabdataset/mov-aad
MOV-AAD: Multimodal Auditory Attention Dataset Official dataset repository for MOV-AAD, a large-scale multimodal dataset designed for investigating selective auditory attention decoding (AAD), spatial audio localization, and cross-modal peripheral physiological tracking during dynamic, naturalistic conversations with moving sound sources. π¬ 10-Minute Walkthrough & Overview βΆοΈ Watch Full 10-Minβ¦ See the full description on the dataset page: https://huggingface.co/datasets/naplabdataset/mov-aad.
MOV-AAD: Multimodal Auditory Attention Dataset
 <br>
  [](https://huggingface.co/datasets/naplabdataset/mov-aad) 
Official dataset repository for MOV-AAD, a large-scale multimodal dataset designed for investigating selective auditory attention decoding (AAD), spatial audio localization, and cross-modal peripheral physiological tracking during dynamic, naturalistic conversations with moving sound sources.
π¬ 10-Minute Walkthrough & Overview
<table style="width: 100%; border: none; border-collapse: collapse;"> <tr style="border: none;"> <td align="center" valign="top" style="width: 50%; border: none; padding: 10px; text-align: center;"> <center> <a href="https://www.youtube.com/watch?v=PbIi4rVktc0" target="blank" title="Watch MOV-AAD Walkthrough on YouTube"> <img src="https://img.youtube.com/vi/PbIi4rVktc0/maxresdefault.jpg" alt="MOV-AAD Video Walkthrough" style="height: 300px; width: 100%; object-fit: contain; border-radius: 8px; border: 1px solid #e1e4e8; box-shadow: 0 2px 8px rgba(0,0,0,0.06); display: block; margin: 0 auto;"> </a> <div style="margin-top: 10px; font-size: 0.9rem; text-align: center;"> βΆοΈ <a href="https://www.youtube.com/watch?v=PbIi4rVktc0" target="blank"><b>Watch Full 10-Min Walkthrough on YouTube</b></a> </div> <div style="margin-top: 4px; font-size: 0.8rem; color: #666; text-align: center;"> (Experimental paradigm, sensor alignment, & demos) </div> </center> </td> <td align="center" valign="top" style="width: 50%; border: none; padding: 10px; text-align: center;"> <center> <a href="https://cdn-uploads.huggingface.co/production/uploads/6abead49900307b51fa24fd2/kZSxV5ZGmZ0XSacamydGk.png" target="blank" title="View Full High-Res Method Vector (PDF)"> <img src="https://cdn-uploads.huggingface.co/production/uploads/6abead49900307b51fa24fd2/kZSxV5ZGmZ0XSacamydGk.png" alt="MOV-AAD Experimental Method & Paradigm" style="height: 300px; width: 100%; object-fit: contain; border-radius: 8px; border: 1px solid #e1e4e8; box-shadow: 0 2px 8px rgba(0,0,0,0.06); display: block; margin: 0 auto;"> </a> <div style="margin-top: 10px; font-size: 0.9rem; text-align: center;"> π <a href="https://cdn-uploads.huggingface.co/production/uploads/6abead49900307b51fa24fd2/kZSxV5ZGmZ0XSacamydGk.png" target="blank"><b>Experimental Paradigm & System Pipeline</b></a> </div> <div style="margin-top: 4px; font-size: 0.8rem; color: #666; text-align: center;"> (Click image or link to open high-res figure) </div> </center> </td> </tr> </table>
π Citation
If you use MOV-AAD in your research, please cite our **Interspeech 2026 paper**:
@inproceedings{he26g_interspeech,
title = {{MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations}},
author = {Xiaomin He and Vishal Choudhari and Tristan J. Spratt and Aarya Raghavan and Richard T. Lee and Nima Mesgarani},
year = {2026},
booktitle = {{Interspeech 2026 [Long Track]}},
pages = {3237--3246},
doi = {10.21437/Interspeech.2026-3556},
issn = {2958-1796}
}π Dataset At A Glance
π οΈ Experimental Paradigms & Data Composition
The dataset encompasses four sequential experimental tasks per participant, transition from structured sensory validation to ecologically valid, naturalistic listening scenarios:
- Repeated Sentence Task (120 trials): A baseline neural reliability check featuring shuffled short sentences repeated without background noise to evaluate within-subject response consistency.
- Localization Task (54 trials): A spatial perception task conducted prior to the main experiment to familiarize participants with spatial cues and assess their localization accuracy.
- Single-Conversation Task (40 trials): An active listening task where participants track naturalistic, continuous conversational narratives from a single dynamically moving sound source.
- Multi-Conversation Task (56 trials): A selective attention task requiring participants to attend to one target conversational stream while ignoring a competing co-spatial distractor.
Parameters Summary Matrix
π¬ Recording Modalities Specification
All signals were synchronously recorded and streamed through g.tec HIamp and a Simulink GUI at 1200 Hz, with a hardware 60 Hz notch filter applied during acquisition.
- EEG: 64 channels, 1200 Hz high-density neural recordings via
g.tec g.HIamp. - Pupil Dilation: Binocular measurements captured at 60 Hz via
Tobii Pro Nanoand upsampled with aligned temporal interpolation to 1200 Hz. - Gaze Location: Screen-coordinate gaze tracking (X and Y axes) mapped to screen size (display dimensions: 52 Γ 32 cm, viewing distance: roughly 60 cm to account for natural posture variation), recorded at 60 Hz and unified to 1200 Hz.
- Respiration Flow: Nasal airflow monitoring.
- Respiration Effort: Thoracic expansion belt (chest) tracking.
- GSR: Galvanic skin response.
- PPG: Raw optical photoplethysmography waveform.
- Heart Rate: Derived beat-by-beat heart rate extracted from raw PPG streams.
- SpOβ: Peripheral oxygen saturation monitoring.
- Temperature: Skin temperature monitoring, recorded from the dorsal surface of the non-dominant hand.
- Accelerometer: Triaxial motion tracking sampled at 1200 Hz, mounted on the chair back to capture gross body and seat vibrations.
π Dataset Directory Structure
The released dataset is organized by experimental task. The four task folders are arranged in parallel, with task-specific recordings and stimuli stored together. For the conversation tasks, each subject recording file contains all trials from that task for the participant, including the synchronized neural, physiological, and behavioral modalities.
MOV-AAD/
βββ SC/ # Single-Conversation Task
β βββ recordings/ # One subject file per participant; all SC trials and synchronized modalities
β β βββ sub-001.*
β β βββ sub-002.*
β β βββ ...
β βββ stimulus/
β βββ audio/ # Trial-level audio presented during SC
β βββ trajectory/ # Trial-level source-motion trajectories
β
βββ MC/ # Multi-Conversation Task
β βββ recordings/ # One subject file per participant; all MC trials and synchronized modalities
β β βββ sub-001.*
β β βββ sub-002.*
β β βββ ...
β βββ stimulus/
β βββ audio/ # Trial-level target/distractor audio used during MC
β βββ trajectory/ # Trial-level source-motion trajectories
β
βββ localization/ # Localization Task
β βββ recordings/ # Subject-level behavioral response recordings
β β βββ sub-001.* # Localization choice responses
β β βββ ...
β βββ stimulus/ # Stimuli used in the localization task
β
βββ repeated_sentence/ # Repeated Sentence Task
β βββ recordings/ # Subject-level recordings for the repeated-sentence task
β β βββ sub-001.*
β β βββ ...
β βββ stimulus/ # Repeated-sentence stimuli
β
βββ preprocessing_scripts/ # Preprocessing and alignment scripts
βββ README.md
βββ dataset_info.jsonOrganization Notes
- Task-first organization: Data are grouped by experimental task rather than by participant.
- SC and MC recordings: Each subject file contains that participant's complete set of trials for the corresponding conversation task, with all available synchronized modalities stored together.
- Conversation stimuli: Audio and source trajectories are stored under each conversation task's
stimulus/directory so that each trial can be paired with the exact presented stimulus and motion path. - Localization task: The released subject recordings contain the behavioral localization responses, corresponding to the participant's multiple-choice spatial reports, together with the task stimuli.
- Repeated Sentence task: Subject recordings and the corresponding repeated-sentence stimuli are stored within the same task-level structure.
π οΈ Preprocessing & Data Alignment Notes
To ensure high data reproducibility, the released dataset provides clean, aligned pipelines processed as follows:
- EEG Artifact Handling & Channel Interpolation
- EEG channels exhibiting abnormal cross-trial variance or high-frequency impedance spikes were automatically flagged.
- Bad channels were reconstructed using Spherical Spline Interpolation (following standard
EEGLABprocedures [Delorme & Makeig, 2004]). - Note: Peripheral modalities are released in a minimally processed form (hardware notch filtering only) to retain raw autonomic dynamics.
- Trial Epoching & Windowing
- Conversation Tasks (SC & MC): Each continuous trial includes a 3-second pre-onset baseline segment prior to speech onset.
- Baseline Period: Vital for calculating trial-level neural entrainment, baseline normalization, or calculating time-frequency relative power.
- Cross-Modal Temporal Realignment
- Primary sync was established via hardware trigger pulses routed simultaneously into the
g.tecdigital input channel. - Fine-grained temporal jitter was eliminated post-hoc via cross-correlation analysis between the recorded acoustic playback loop and the original master stimulus waveforms, ensuring sub-millisecond precision.
π» Download & Access
Because of the massive scale of the multi-channel recordings, the full dataset archive is hosted externally across dedicated data repositories.
π¦ Access Channels
- Google Drive (Primary Archive β Controlled Access)
- Link: Google Drive Master Archive
- Access Note: Click Request Access on Google Drive. Access requests are approved manually in accordance with our institutional governance protocol. You can download the complete archive or select specific task folders (
SC/,MC/,localization/,repeated_sentence/) as needed.
- Hugging Face Hub (In Preparation)
- Link: `naplabdataset/mov-aad`
- Access Note: Dataset upload is currently in progress. Direct CLI and programmatic loading via Hugging Face will be supported upon completion.
