Team Ai
Datasetpublic

naplabdataset/mov-aad

MOV-AAD: Multimodal Auditory Attention Dataset Official dataset repository for MOV-AAD, a large-scale multimodal dataset designed for investigating selective auditory attention decoding (AAD), spatial audio localization, and cross-modal peripheral physiological tracking during dynamic, naturalistic conversations with moving sound sources. 🎬 10-Minute Walkthrough & Overview ▢️ Watch Full 10-Min… See the full description on the dataset page: https://huggingface.co/datasets/naplabdataset/mov-aad.

sourceHugging Facecc-by-4.0updated 5d agoView on Hugging Face
1likes195downloads
Dataset Card

MOV-AAD: Multimodal Auditory Attention Dataset

![Award](https://interspeech2026.org) <br>

![Paper](https://www.isca-archive.org/interspeech2026/he26ginterspeech.html) ![Google Drive](https://drive.google.com/open?id=19D0o1WT7R7lDVCC-l7-WhYWsGXVYqpaS&usp=drivefs) [![Hugging Face](https://img.shields.io/badge/πŸ€—HuggingFace-Public_Dataset-FFD21E.svg)](https://huggingface.co/datasets/naplabdataset/mov-aad) ![License](https://creativecommons.org/licenses/by/4.0/)

Official dataset repository for MOV-AAD, a large-scale multimodal dataset designed for investigating selective auditory attention decoding (AAD), spatial audio localization, and cross-modal peripheral physiological tracking during dynamic, naturalistic conversations with moving sound sources.

🎬 10-Minute Walkthrough & Overview

<table style="width: 100%; border: none; border-collapse: collapse;"> <tr style="border: none;"> <td align="center" valign="top" style="width: 50%; border: none; padding: 10px; text-align: center;"> <center> <a href="https://www.youtube.com/watch?v=PbIi4rVktc0" target="blank" title="Watch MOV-AAD Walkthrough on YouTube"> <img src="https://img.youtube.com/vi/PbIi4rVktc0/maxresdefault.jpg" alt="MOV-AAD Video Walkthrough" style="height: 300px; width: 100%; object-fit: contain; border-radius: 8px; border: 1px solid #e1e4e8; box-shadow: 0 2px 8px rgba(0,0,0,0.06); display: block; margin: 0 auto;"> </a> <div style="margin-top: 10px; font-size: 0.9rem; text-align: center;"> ▢️ <a href="https://www.youtube.com/watch?v=PbIi4rVktc0" target="blank"><b>Watch Full 10-Min Walkthrough on YouTube</b></a> </div> <div style="margin-top: 4px; font-size: 0.8rem; color: #666; text-align: center;"> (Experimental paradigm, sensor alignment, & demos) </div> </center> </td> <td align="center" valign="top" style="width: 50%; border: none; padding: 10px; text-align: center;"> <center> <a href="https://cdn-uploads.huggingface.co/production/uploads/6abead49900307b51fa24fd2/kZSxV5ZGmZ0XSacamydGk.png" target="blank" title="View Full High-Res Method Vector (PDF)"> <img src="https://cdn-uploads.huggingface.co/production/uploads/6abead49900307b51fa24fd2/kZSxV5ZGmZ0XSacamydGk.png" alt="MOV-AAD Experimental Method & Paradigm" style="height: 300px; width: 100%; object-fit: contain; border-radius: 8px; border: 1px solid #e1e4e8; box-shadow: 0 2px 8px rgba(0,0,0,0.06); display: block; margin: 0 auto;"> </a> <div style="margin-top: 10px; font-size: 0.9rem; text-align: center;"> πŸ” <a href="https://cdn-uploads.huggingface.co/production/uploads/6abead49900307b51fa24fd2/kZSxV5ZGmZ0XSacamydGk.png" target="blank"><b>Experimental Paradigm & System Pipeline</b></a> </div> <div style="margin-top: 4px; font-size: 0.8rem; color: #666; text-align: center;"> (Click image or link to open high-res figure) </div> </center> </td> </tr> </table>

πŸ“œ Citation

If you use MOV-AAD in your research, please cite our **Interspeech 2026 paper**:

bibtex
@inproceedings{he26g_interspeech,
  title     = {{MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations}},
  author    = {Xiaomin He and Vishal Choudhari and Tristan J. Spratt and Aarya Raghavan and Richard T. Lee and Nima Mesgarani},
  year      = {2026},
  booktitle = {{Interspeech 2026 [Long Track]}},
  pages     = {3237--3246},
  doi       = {10.21437/Interspeech.2026-3556},
  issn      = {2958-1796}
}

πŸ“Š Dataset At A Glance

**10** Synchronized Modalities**1200 Hz** Unified Sampling Rate**50** Healthy Subjects
64-ch EEG, Gaze, Pupil, Respiration Airflow, Respiration Effort, GSR, PPG, SpOβ‚‚, Temperature, MotionFully aligned across all neural & physiological data streamsAge 24 Β± 4.5 years, verified normal hearing status
**~75 Min** Total Duration**4** Distinct Tasks**Rich** Behavioral Tracking
Massive continuous recording per participant for AAD modelingFrom structured validation to naturalistic conversationsActive repeated-word detection & precise localization reports

πŸ› οΈ Experimental Paradigms & Data Composition

The dataset encompasses four sequential experimental tasks per participant, transition from structured sensory validation to ecologically valid, naturalistic listening scenarios:

  1. 1.Repeated Sentence Task (120 trials): A baseline neural reliability check featuring shuffled short sentences repeated without background noise to evaluate within-subject response consistency.
  2. 2.Localization Task (54 trials): A spatial perception task conducted prior to the main experiment to familiarize participants with spatial cues and assess their localization accuracy.
  3. 3.Single-Conversation Task (40 trials): An active listening task where participants track naturalistic, continuous conversational narratives from a single dynamically moving sound source.
  4. 4.Multi-Conversation Task (56 trials): A selective attention task requiring participants to attend to one target conversational stream while ignoring a competing co-spatial distractor.

Parameters Summary Matrix

Task / ConditionExperimental PurposeStimuli & SpeakersSpatial ConfigurationBackground NoiseTotal Stimuli / DurationAvailable Modalities
Repeated Sentence TaskSplit-half & odd-even reliability check6 short sentences (shuffled); 3M/3F talkersDiotic presentation (None)None120 trials<br>(~12 min)βœ“ EEG & Physio<br>βœ“ Behavioral
Localization TaskFamiliarize spatial cues & assess perception accuracyIndependent short sentences; Single talker/trialStatic; 9 coordinates (-90Β° to +90Β° in 22.5Β° steps)None54 sentences<br>(Short segments)- No Neural<br>βœ“ Behavior Metrics
Single-Conversation TaskTrack neural tracking of speech with spatial changeContinuous dialogue context; Multi-turn talkersHRTF-based dynamic moving (-90Β° to +90Β°); RMS-matchedDiotic pedestrian/babble noise (-9, -12 dB)40 trials<br>(~30 min)βœ“ EEG & Physio<br>βœ“ Behavior & Audio
Multi-Conversation TaskEvaluate selective Auditory Attention Decoding (AAD)Two parallel stories; Continuous context & turn-takingTwo independent HRTF sources moving dynamically within Β±90Β°Diotic pedestrian/babble noise (-9, -12 dB)56 trials<br>(~45 min)βœ“ EEG & Physio<br>βœ“ Behavior & Audio

πŸ”¬ Recording Modalities Specification

All signals were synchronously recorded and streamed through g.tec HIamp and a Simulink GUI at 1200 Hz, with a hardware 60 Hz notch filter applied during acquisition.

  • β€”EEG: 64 channels, 1200 Hz high-density neural recordings via g.tec g.HIamp.
  • β€”Pupil Dilation: Binocular measurements captured at 60 Hz via Tobii Pro Nano and upsampled with aligned temporal interpolation to 1200 Hz.
  • β€”Gaze Location: Screen-coordinate gaze tracking (X and Y axes) mapped to screen size (display dimensions: 52 Γ— 32 cm, viewing distance: roughly 60 cm to account for natural posture variation), recorded at 60 Hz and unified to 1200 Hz.
  • β€”Respiration Flow: Nasal airflow monitoring.
  • β€”Respiration Effort: Thoracic expansion belt (chest) tracking.
  • β€”GSR: Galvanic skin response.
  • β€”PPG: Raw optical photoplethysmography waveform.
  • β€”Heart Rate: Derived beat-by-beat heart rate extracted from raw PPG streams.
  • β€”SpOβ‚‚: Peripheral oxygen saturation monitoring.
  • β€”Temperature: Skin temperature monitoring, recorded from the dorsal surface of the non-dominant hand.
  • β€”Accelerometer: Triaxial motion tracking sampled at 1200 Hz, mounted on the chair back to capture gross body and seat vibrations.

πŸ“‚ Dataset Directory Structure

The released dataset is organized by experimental task. The four task folders are arranged in parallel, with task-specific recordings and stimuli stored together. For the conversation tasks, each subject recording file contains all trials from that task for the participant, including the synchronized neural, physiological, and behavioral modalities.

directory
MOV-AAD/
β”œβ”€β”€ SC/                              # Single-Conversation Task
β”‚   β”œβ”€β”€ recordings/                  # One subject file per participant; all SC trials and synchronized modalities
β”‚   β”‚   β”œβ”€β”€ sub-001.*
β”‚   β”‚   β”œβ”€β”€ sub-002.*
β”‚   β”‚   └── ...
β”‚   └── stimulus/
β”‚       β”œβ”€β”€ audio/                   # Trial-level audio presented during SC
β”‚       └── trajectory/              # Trial-level source-motion trajectories
β”‚
β”œβ”€β”€ MC/                              # Multi-Conversation Task
β”‚   β”œβ”€β”€ recordings/                  # One subject file per participant; all MC trials and synchronized modalities
β”‚   β”‚   β”œβ”€β”€ sub-001.*
β”‚   β”‚   β”œβ”€β”€ sub-002.*
β”‚   β”‚   └── ...
β”‚   └── stimulus/
β”‚       β”œβ”€β”€ audio/                   # Trial-level target/distractor audio used during MC
β”‚       └── trajectory/              # Trial-level source-motion trajectories
β”‚
β”œβ”€β”€ localization/                    # Localization Task
β”‚   β”œβ”€β”€ recordings/                  # Subject-level behavioral response recordings
β”‚   β”‚   β”œβ”€β”€ sub-001.*                # Localization choice responses
β”‚   β”‚   └── ...
β”‚   └── stimulus/                    # Stimuli used in the localization task
β”‚
β”œβ”€β”€ repeated_sentence/               # Repeated Sentence Task
β”‚   β”œβ”€β”€ recordings/                  # Subject-level recordings for the repeated-sentence task
β”‚   β”‚   β”œβ”€β”€ sub-001.*
β”‚   β”‚   └── ...
β”‚   └── stimulus/                    # Repeated-sentence stimuli
β”‚
β”œβ”€β”€ preprocessing_scripts/           # Preprocessing and alignment scripts
β”œβ”€β”€ README.md
└── dataset_info.json

Organization Notes

  • β€”Task-first organization: Data are grouped by experimental task rather than by participant.
  • β€”SC and MC recordings: Each subject file contains that participant's complete set of trials for the corresponding conversation task, with all available synchronized modalities stored together.
  • β€”Conversation stimuli: Audio and source trajectories are stored under each conversation task's stimulus/ directory so that each trial can be paired with the exact presented stimulus and motion path.
  • β€”Localization task: The released subject recordings contain the behavioral localization responses, corresponding to the participant's multiple-choice spatial reports, together with the task stimuli.
  • β€”Repeated Sentence task: Subject recordings and the corresponding repeated-sentence stimuli are stored within the same task-level structure.

πŸ› οΈ Preprocessing & Data Alignment Notes

To ensure high data reproducibility, the released dataset provides clean, aligned pipelines processed as follows:

  • β€”EEG Artifact Handling & Channel Interpolation
  • β€”EEG channels exhibiting abnormal cross-trial variance or high-frequency impedance spikes were automatically flagged.
  • β€”Bad channels were reconstructed using Spherical Spline Interpolation (following standard EEGLAB procedures [Delorme & Makeig, 2004]).
  • β€”Note: Peripheral modalities are released in a minimally processed form (hardware notch filtering only) to retain raw autonomic dynamics.
  • β€”Trial Epoching & Windowing
  • β€”Conversation Tasks (SC & MC): Each continuous trial includes a 3-second pre-onset baseline segment prior to speech onset.
  • β€”Baseline Period: Vital for calculating trial-level neural entrainment, baseline normalization, or calculating time-frequency relative power.
  • β€”Cross-Modal Temporal Realignment
  • β€”Primary sync was established via hardware trigger pulses routed simultaneously into the g.tec digital input channel.
  • β€”Fine-grained temporal jitter was eliminated post-hoc via cross-correlation analysis between the recorded acoustic playback loop and the original master stimulus waveforms, ensuring sub-millisecond precision.

πŸ’» Download & Access

Because of the massive scale of the multi-channel recordings, the full dataset archive is hosted externally across dedicated data repositories.

πŸ“¦ Access Channels

  • β€”Google Drive (Primary Archive β€” Controlled Access)
  • β€”Link: Google Drive Master Archive
  • β€”Access Note: Click Request Access on Google Drive. Access requests are approved manually in accordance with our institutional governance protocol. You can download the complete archive or select specific task folders (SC/, MC/, localization/, repeated_sentence/) as needed.
  • β€”Hugging Face Hub (In Preparation)
  • β€”Link: `naplabdataset/mov-aad`
  • β€”Access Note: Dataset upload is currently in progress. Direct CLI and programmatic loading via Hugging Face will be supported upon completion.