datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpatialAudio
SpatialAudio
This repo hosts the dataset and models of "BAT: Learning to Reason about Spatial Sounds with Large Language Models" [ICML 2024 bib].
Spatial Audio Dataset (Mono/Binaural/Ambisonics)
AudioSet (Anechoic Audio Source)
We provide Balanced train and Evaluation set for your convenience. You can download from SpatialAudio.
For the Unbalanced train set, please refer to Official AudioSet.
Metadata can be downloaded from metadata.
AudioSet
├──… See the full description on the dataset page: https://huggingface.co/datasets/zhisheng01/SpatialAudio.gram-scenes-as2m-16k
GRAM-style binaural scenes from AudioSet-2M, 16 kHz
AudioSet-2M rendered to binaural stereo following the scene-synthesis recipe of
GRAM (arXiv:2506.00934), used to pretrain stereo BEATs encoders.
64 shard tars, 829.5 GB, 2,020,706 clips — one scene
per clip of AudioSet's unbalanced_train set.
Rendering
Per clip, as in GRAM-T (dataset_functions.py): AudioSet mono -> RMS -14 dBFS ->
10 s pad/truncate -> convolve with a binaural RIR; noise from WHAM (tr only)… See the full description on the dataset page: https://huggingface.co/datasets/spatial-audio-learning/gram-scenes-as2m-16k.SpatialAudio
SpatialAudio
This repo hosts the dataset and models of "BAT: Learning to Reason about Spatial Sounds with Large Language Models" [ICML 2024 bib].
Spatial Audio Dataset (Mono/Binaural/Ambisonics)
AudioSet (Anechoic Audio Source)
We provide Balanced train and Evaluation set for your convenience. You can download from SpatialAudio.
For the Unbalanced train set, please refer to Official AudioSet.
Metadata can be downloaded from metadata.
AudioSet
├──… See the full description on the dataset page: https://huggingface.co/datasets/yongr/SpatialAudio.SpatialAudioSpatialAudio
