Team Ai
Datasetpublic

davidggphy/librispeech-arpabet-processed

LibriSpeech ARPAbet Processed Dataset Pre-processed dataset for training ARPAbet phoneme recognition models using CTC loss. Dataset Description This dataset is derived from LibriSpeech (train-clean-100 split) with the following preprocessing: Audio: Resampled to 16kHz, normalized using Wav2Vec2 feature extractor Labels: Text transcriptions converted to ARPAbet phoneme sequences using CMU Pronouncing Dictionary Filtering: Samples with out-of-vocabulary words (not… See the full description on the dataset page: https://huggingface.co/datasets/davidggphy/librispeech-arpabet-processed.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes185downloads

davidggphy/librispeech-arpabet-processed · main · files are served by the source, never re-hosted here