Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes32 downloads5mo agoHugging Face02bingbangboom /cleaned-asr-transcriptstexttext-generation10K<n<100K1 likes29 downloads7mo agoHugging Face03exoarbuus /goatis-transcripts Goatis / Sv3rige Video Transcripts Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the current Goatis channel (2019–2026). This is the dataset behind goatis.net, a searchable archive in the style of aajonus.net. What makes it more than raw ASR Every video was processed with speaker identification, not just transcription. He mostly reacts to other people's videos, so a… See the full description on the dataset page: https://huggingface.co/datasets/exoarbuus/goatis-transcripts.textautomatic-speech-recognition1K<n<10K0 likes11 downloads3mo agoHugging Face04SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.