Team Ai
Datasetpublic

confit/librispeech-sid-parquet

LibriSpeech Speaker Identification LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. However, although LibriSpeech is very popular in ASR tasks, we use LibriSpeech database as a speaker identification task. We follow SincNet paper official split for training and… See the full description on the dataset page: https://huggingface.co/datasets/confit/librispeech-sid-parquet.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes225downloads
Dataset Card

LibriSpeech Speaker Identification

LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.

However, although LibriSpeech is very popular in ASR tasks, we use LibriSpeech database as a speaker identification task. We follow SincNet paper official split for training and evaluation. It has 2484 number of classes (unique speakers) with a total of 21933 (14481 / 7452) samples.

Citation

bibtex
@misc{ravanelli2019speaker,
  title={Speaker Recognition from Raw Waveform with SincNet}, 
  author={Mirco Ravanelli and Yoshua Bengio},
  year={2019},
  eprint={1808.00158},
  archivePrefix={arXiv},
  primaryClass={eess.AS}
}
bibtex
@misc{ravanelli2019speaker,
  title={Speaker Recognition from Raw Waveform with SincNet}, 
  author={Mirco Ravanelli and Yoshua Bengio},
  year={2019},
  eprint={1808.00158},
  archivePrefix={arXiv},
  primaryClass={eess.AS}
}