Zackh/expresso-contextual
The Expresso Dataset [paper] [demo samples] [Original repository] Introduction The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). Improvised Dialogues This HuggingFace dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Zackh/expresso-contextual.
The Expresso Dataset
[[paper]](https://arxiv.org/abs/2308.05725) [[demo samples]](https://speechbot.github.io/expresso/) [[Original repository]](https://github.com/facebookresearch/textlesslib/tree/main/examples/expresso/dataset)
Introduction
The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised).
Improvised Dialogues
This HuggingFace dataset contains only the improvised dialogues. The format is two-channel audio of entire conversations and json of the turns in the conversation.
The json representing the turns for one conversation has the following structure:
[
{
"speaker": "ex02",
"start_time_ms": 0.0,
"end_time_ms": 23760.0,
"channel": 1
"text": "Um, hi my name is Jackie or Jacqueline, hi. Um, I am from Toronto Ontario, Canada. I was raised there up until I decided to move here to the US in 2013 and I moved to LA. Um, yeah and I've been in Hollywood ever since, um.",
},
{
"speaker": ...
},
...
]Read Speech
See ylacombe/expresso for the Expresso's read speech examples.
Audio Quality
The audio was recorded in a professional recording studio with minimal background noise at 48kHz/24bit. The files for read speech and singing are in a mono wav format; and for the dialog section in stereo (one channel per actor), where the original flow of turn-taking is preserved.
License
The Expresso dataset is distributed under the CC BY-NC 4.0 license.
Reference
For more information, see the paper: EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis, Tu Anh Nguyen, Wei-Ning Hsu, Antony D'Avirro, Bowen Shi, Itai Gat, Maryam Fazel-Zarani, Tal Remez, Jade Copet, Gabriel Synnaeve, Michael Hassid, Felix Kreuk, Yossi Adi⁺, Emmanuel Dupoux⁺, INTERSPEECH 2023.
