Team Ai
Datasetpublic

Zackh/expresso-contextual

The Expresso Dataset [paper] [demo samples] [Original repository] Introduction The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). Improvised Dialogues This HuggingFace dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Zackh/expresso-contextual.

sourceHugging Facecc-by-nc-4.0updated 4mo agoView on Hugging Face
3likes110downloads
Dataset Card

The Expresso Dataset

[[paper]](https://arxiv.org/abs/2308.05725) [[demo samples]](https://speechbot.github.io/expresso/) [[Original repository]](https://github.com/facebookresearch/textlesslib/tree/main/examples/expresso/dataset)

Introduction

The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised).

Improvised Dialogues

This HuggingFace dataset contains only the improvised dialogues. The format is two-channel audio of entire conversations and json of the turns in the conversation.

The json representing the turns for one conversation has the following structure:

json
[
  {
    "speaker": "ex02",
    "start_time_ms": 0.0,
    "end_time_ms": 23760.0,
    "channel": 1
    "text": "Um, hi my name is Jackie or Jacqueline, hi. Um, I am from Toronto Ontario, Canada. I was raised there up until I decided to move here to the US in 2013 and I moved to LA. Um, yeah and I've been in Hollywood ever since, um.",
  },
  {
    "speaker": ...
  },
  ...
]

Read Speech

See ylacombe/expresso for the Expresso's read speech examples.

Audio Quality

The audio was recorded in a professional recording studio with minimal background noise at 48kHz/24bit. The files for read speech and singing are in a mono wav format; and for the dialog section in stereo (one channel per actor), where the original flow of turn-taking is preserved.

License

The Expresso dataset is distributed under the CC BY-NC 4.0 license.

Reference

For more information, see the paper: EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis, Tu Anh Nguyen, Wei-Ning Hsu, Antony D'Avirro, Bowen Shi, Itai Gat, Maryam Fazel-Zarani, Tal Remez, Jade Copet, Gabriel Synnaeve, Michael Hassid, Felix Kreuk, Yossi Adi⁺, Emmanuel Dupoux⁺, INTERSPEECH 2023.