Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anuj-inavlabs /Thinkspark-v2-270m-training-data ThinkSpark-v2-350M — training data Full-duplex floor-controller (Section 8) training corpus: playable audio + text, paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain, gender, prosody, agent text) and Soniox character-level timestamps. Dataset Viewer Default split is parquet with a real Audio feature — a player renders inline next to the text in the Hub UI: column type description audio Audio playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.text-to-speech1K<n<10K0 likes1.7k downloads1mo agoHugging Face02KitTzk /lao_stt_training_data Lao Speech-to-Text Training Data ຊຸດຂໍ້ມູນນີ້ຖືກຈັດກຽມຂຶ້ນມາເພື່ອໃຊ້ສຳລັບການເທຣນ ແລະ ປັບແຕ່ງ (Fine-tuning) ໂມເດວ Speech-to-Text (ເຊັ່ນ OpenAI Whisper) ສຳລັບພາສາລາວ. ໂຄງສ້າງຂອງຂໍ້ມູນ (Dataset Structure) Train set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ train/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ train.csv Validation set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ validation/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ validation.csv ຮູບແບບຂໍ້ມູນໃນໄຟລ໌ CSV: audio: ເສັ້ນທາງໄປຫາໄຟລ໌ສຽງ (e.g., train/audio25000.wav)… See the full description on the dataset page: https://huggingface.co/datasets/KitTzk/lao_stt_training_data.audioautomatic-speech-recognition1K<n<10K1 likes163 downloads4mo agoHugging Face03nlpctx /tts-training-dataset Human Reviewed Telugu-English TTS Dataset A manually reviewed multilingual TTS dataset created from publicly available educational and speech content. Dataset Splits & Distribution Metrics balanced_60min Split Total Segments: 120 Total Duration: 60.00 minutes Unique Speakers: 3 Distribution Breakdowns: Language Distribution: en-IN: 60 segments (30.00 minutes) te-IN: 60 segments (30.00 minutes) Style Distribution: analytical: 27… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tts-training-dataset.audiotext-to-speechn<1K0 likes31 downloads4mo agoHugging Face04SRP-base-model-training /kazakh_speech_dataset_ksdgatedKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api. Dataset info: 813 Speakers with 500 samples for 4 speakers with 250 samples for 809 speakers Male/female 555 Hours Guides Load data 1 Replace the export HF_HOME with your HF_HOME path from datasets import load_dataset # export HF_HOME="/data/vladimir_albrekht/hf_cache" ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.audioautomatic-speech-recognition100K<n<1M2 likes30 downloads1y agoHugging Face05DigiGreen /KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers. Process of collection of data: Selected users were given the option of doing a task and getting paid for it. The users were supposed to record the sentence as it appeared on the screen. The audio file thus obtained was validated matched with the sentences to fine tune the model. Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/KikuyuASR_trainingdataset.automatic-speech-recognition10K<n<100K0 likes27 downloads2y agoHugging Face06CGIAR /KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers. Process of collection of data: Selected users were given the option of doing a task and getting paid for it. The users were supposed to record the sentence as it appeared on the screen. The audio file thus obtained was validated matched with the sentences to fine tune the model. Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/KikuyuASR_trainingdataset.automatic-speech-recognition10K<n<100K0 likes24 downloads2y agoHugging Face07jo-05 /KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers. Process of collection of data: Selected users were given the option of doing a task and getting paid for it. The users were supposed to record the sentence as it appeared on the screen. The audio file thus obtained was validated matched with the sentences to fine tune the model. Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/jo-05/KikuyuASR_trainingdataset.automatic-speech-recognition10K<n<100K0 likes14 downloads7mo agoHugging Face08OpenLLM-France /Luciole-Audio-Training-Dataset Luciole Audio Training Dataset Dataset description Luciole Audio Training Dataset is a large, multilingual, multi-task collection of audio–text conversations used to train the OpenLLM-France Luciole audio-language models. It adapts a text LLM to understand audio by pairing speech, music and environmental sounds with instruction-style dialogues (transcription, translation, spoken question answering, audio/music/sound captioning and question answering, speaker and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Audio-Training-Dataset.textautomatic-speech-recognition10M<n<100M1 likes6 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.