KimZoey/reaction_data
🧠 Overview Reaction Data v1.0 is a curated multimodal dataset built from YouTube human reaction videos, designed for studying singing voice synthesis (SVS) evaluation, audio–text alignment, and natural-language reaction modeling.Each included creator (e.g., Alex Hefner, Beth Roars, The Charismatic Voice) provides a distinct commentary style and persona, offering diverse reactions across genres and vocal styles. Each reviewer’s folder contains processed data, transcripts, and… See the full description on the dataset page: https://huggingface.co/datasets/KimZoey/reaction_data.
🧠 Overview
Reaction Data v1.0 is a curated multimodal dataset built from YouTube human reaction videos, designed for studying singing voice synthesis (SVS) evaluation, audio–text alignment, and natural-language reaction modeling. Each included creator (e.g., Alex Hefner, Beth Roars, The Charismatic Voice) provides a distinct commentary style and persona, offering diverse reactions across genres and vocal styles.
Each reviewer’s folder contains processed data, transcripts, and raw video material, enabling flexible multimodal experiments.
📁 Dataset Structure
Each reviewer (e.g., Alex_Hefner, BethRoars, The_Charismatic_Voice, etc.) has three corresponding archives:
Additionally, the file user_info.json stores reviewer metadata and file list.
⚙️ Data Processing Pipeline
The dataset is created through the following pipeline:
- Video & Subtitle Retrieval
- Download with
yt-dlp. - Retrieve subtitles via the YouTube Transcript API.
- Audio Extraction & Diarization
- Extract audio from each video.
- Apply pyannote.audio for speaker diarization and merge utterances per speaker.
- Audio Classification
- Each utterance is classified as singing or spoken commentary using
MIT/ast-finetuned-audioset-10-10-0.4593.
- Segment Alignment
- Align contiguous singing segments with subsequent critic commentary.
- Match commentary timestamps with subtitles to extract review text.
- Metadata Integration
- Attach reviewer persona metadata (from channel introductions).
- Add song metadata (from Wikipedia).
- Form complete multimodal training samples.
- Quality Filtering
- Remove samples with:
- Empty subtitles or missing text
- Audio shorter than 10 seconds
- Text shorter than 8 words → Ensures linguistically and acoustically rich data.
💡 Research Motivation
This dataset supports research on:
- Generating and evaluating natural-language feedback for singing performances
- Modeling human-like reaction patterns across multiple reviewer personas
- Assessing singing voice synthesis systems under realistic multimodal conditions
By including authentic reaction speech, expressive commentary, and varied recording setups, it fosters robust model generalization.
🧩 Notes & Limitations
Alternative pipelines (ASR-first, diarization-only) were tested but found less reliable due to:
- Overlapping singing and commentary causing poor segmentation
- Garbled ASR transcriptions under mixed speech
- Subtitle–speaker mismatch for short utterances
The adopted hybrid pipeline provides the best balance between accuracy, diversity, and robustness.
📜 Metadata
- Reviewer list and metadata: stored in
user_info.json - Fields include:
channel_namepersona_descriptionvideo_countlanguageestimated_total_duration
🏷️ Tags
audio-text, multimodal, reaction, singing, music-evaluation, svs, huggingface-datasets, LLM4Music, commentary
