Team Ai
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LEMAS-Project /LEMAS-Dataset-train Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.texttext-to-speech100M<n<1B90 likes6.5k downloads6mo agoHugging Face02danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes120 downloads11mo agoHugging Face03mesolitica /pseudolabel-malaya-speech-stt-train-whisper-large-v3tabularautomatic-speech-recognition1M<n<10M1 likes111 downloads3y agoHugging Face04rustam1221 /uzbek-asr-train-manifests Uzbek ASR Training Manifests The exact training, validation and test splits behind rustam1221/uzbek-asr-gigaam: 974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized, and split by speaker. No audio is copied. Each row is a pointer — a parquet file plus a row index in the upstream dataset — and the training dataloader decodes the audio when the batch is built. That keeps the whole corpus definition at 200 MB instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.textautomatic-speech-recognition1K<n<10K0 likes57 downloads1mo agoHugging Face05theblackcat102 /quantized-librispeech-train-360textautomatic-speech-recognition100K<n<1M0 likes12 downloads3y agoHugging Face06OpenLLM-France /Luciole-Audio-Training-Dataset Luciole Audio Training Dataset Dataset description Luciole Audio Training Dataset is a large, multilingual, multi-task collection of audio–text conversations used to train the OpenLLM-France Luciole audio-language models. It adapts a text LLM to understand audio by pairing speech, music and environmental sounds with instruction-style dialogues (transcription, translation, spoken question answering, audio/music/sound captioning and question answering, speaker and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Audio-Training-Dataset.textautomatic-speech-recognition10M<n<100M1 likes6 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.