Team Ai
Datasetpublic

philgzl/ears

EARS: Expressive Anechoic Recordings of Speech This is a mirror of the Expressive Anechoic Recordings of Speech (EARS) dataset. The original files were converted from WAV to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: Train: 92 hours, 15939 utterances, speakers p001 to p099 Validation: 2 hours, 322 utterances, speakers p100 and p101 Test: 6 hours, 966 utterances, speakers p102 to p107 License: CC BY-NC 4.0 Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/ears.

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
0likes892downloads
Dataset Card

EARS: Expressive Anechoic Recordings of Speech

This is a mirror of the Expressive Anechoic Recordings of Speech (EARS) dataset. The original files were converted from WAV to Opus to reduce the size and accelerate streaming.

Usage

python
import io

import soundfile as sf
from datasets import Features, Value, load_dataset

for item in load_dataset(
    "philgzl/ears",
    split="train",
    streaming=True,
    features=Features({"audio": Value("binary"), "name": Value("string")}),
):
    print(item["name"])
    buffer = io.BytesIO(item["audio"])
    x, fs = sf.read(buffer)
    # do stuff...

Citation

bibtex
@inproceedings{richter2024ears,
  title = {{EARS}: {An} anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation},
  author = {Richter, Julius and Wu, Yi-Chiao and Krenn, Steven and Welker, Simon and Lay, Bunlong and Watanabe, Shinjii and Richard, Alexander and Gerkmann, Timo},
  booktitle = {Proc. Interspeech},
  pages = {4873--4877},
  year = {2024},
}