datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wav2vec2_common_voice_accents_3sada-train-wav2vec2-xls-r-300m-ar-preprocessedwav2vec2-large-xls-r-300m-750-100-onelabelwav2vec2-test-datasetwav2vec2-ru-IIIsimplified_google_speech_commands_wav2vec2_960hwav2vec2-ksponspeech-train2Wav2Vec2_ASVSpoof5_FULLMSPP_WAV2Vec2wav2vec2_common_voice_accentsMSPP_Wav2Vec2_V2wav2vec2-base-lj-demo-colabwav2vec2-large-xls-r-300m-tr-colabwav2vec2_basespot_data_allwav2vec2-ru-IIsimplified-google-speech-commands-wav2vec2-960hwav2vec2-ru-IVwav2vec2-russianPMEmo2019-audio-wav2vec2-processed-valence-trainwav2vec2_phoneme_spot_data_allNMSQA_features_wav2vec2-large-lv60wav2vec2-vd-bird-sound-classification-datasetwav2vec2-ksponspeech-trainprocessed-LA-Wav2Vec2PMEmo2019-audio-wav2vec2-processed-valence-new-regressor-train-datasetwav2vec2-ru-Ienenlhet-wav2vec2-dataset
Enenlhet Wav2Vec2 Dataset
This dataset contains preprocessed audio features and tokenized text for training Wav2Vec2 models on the Enenlhet language.
Dataset Summary
Train: 3,053 examples
Test: 170 examples
Validation: 170 examples
Total: 3,393 examples
Features
input_values: Preprocessed audio features (16kHz, normalized float32 arrays)
labels: Tokenized text as integer sequences
Usage
from datasets import load_dataset
# Load the dataset… See the full description on the dataset page: https://huggingface.co/datasets/sjhuskey/enenlhet-wav2vec2-dataset.sada-validation-wav2vec2-xls-r-300m-ar-preprocessedwav2vec2_processed_spotify-nonstratifieddstc2_audios_input_wav2vec2
