Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes44k downloads2y agoHugging Face02japanese-asr /whisper_transcriptions.mls.wer_10.0audio1M<n<10M2 likes14k downloads2y agoHugging Face03japanese-asr /whisper_transcriptions.reazonspeech.allaudio10M<n<100M4 likes6.2k downloads2y agoHugging Face04japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0audio1M<n<10M3 likes3.3k downloads2y agoHugging Face05Whispering-GPT /linustechtips-transcript-audio Dataset Card for "linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.audioautomatic-speech-recognitionn<1K4 likes2.8k downloads4y agoHugging Face06Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads4y agoHugging Face07nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B10 likes2.2k downloads2y agoHugging Face08japanese-asr /whisper_transcriptions.mlsaudio10M<n<100M1 likes1.2k downloads2y agoHugging Face09ghanaopenai /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes1.2k downloads1mo agoHugging Face10jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face11openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes750 downloads7mo agoHugging Face12japanese-asr /whisper_transcriptions.reazon_speech_all.wer_10.0audio1M<n<10M0 likes651 downloads2y agoHugging Face13japanese-asr /whisper_transcriptions.reazonspeech.largeaudio1M<n<10M0 likes621 downloads3y agoHugging Face14ghananlpcommunity /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes572 downloads1mo agoHugging Face15japanese-asr /whisper_transcriptions.reazonspeech.mediumaudio100K<n<1M0 likes498 downloads3y agoHugging Face16yuriyvnv /synthetic_transcript_pt Portuguese Speech Dataset with Multiple Training Configurations A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms. 🎯 Dataset Configurations Overview This dataset provides three carefully curated subsets to enable comprehensive speech recognition research: Configuration Training Data Validation Test Total Samples Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.audioautomatic-speech-recognition100K<n<1M0 likes434 downloads5mo agoHugging Face17united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes433 downloads7mo agoHugging Face18octava /indonesian-voice-transcription-1.4.9a-raudio10K<n<100K0 likes384 downloads2y agoHugging Face19japanese-asr /whisper_transcriptions.reazonspeech.large.wer_10.0audio1M<n<10M0 likes381 downloads3y agoHugging Face20modulate /entity-transcription-benchmark Entity Transcription Benchmark Measures whether a speech recognition system transcribes named entities correctly — as distinct from word error rate. WER weights every token equally. The tokens that matter for redaction, lookup, routing and search are proper nouns, and they are a small fraction of any transcript. A system can improve WER while getting worse at exactly the words a downstream consumer needs, and nothing in the standard evaluation will show it. 2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.audioautomatic-speech-recognition1K<n<10K4 likes372 downloads26d agoHugging Face21octava /indonesian-voice-transcription-1.4.9a.2audio10K<n<100K0 likes314 downloads2y agoHugging Face22Whispering-GPT /yannick-kilcher-transcript-audio Dataset Card for "yannic-kilcher-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher. Data Fields The dataset is composed by: id: Id of the youtube video. channel:… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannick-kilcher-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes310 downloads4y agoHugging Face23diarizers-community /ami_ihm_with_transcriptsaudion<1K0 likes292 downloads2y agoHugging Face24surry-hills-druid /noagenda-transcripts noagenda transcripts This is the dataset for transcripts of the noagendashow.net podcast. The transcripts are in the data folder. It also contains the source code for generating the transcripts, and the source code for the noagenda-transcripts.net website which searches the transcripts and plays the audio clips for each search result and takes you to the location in the transcript. The code in the repo consists of 2 main parts: A Go CLI for transcribing the audio, creating the… See the full description on the dataset page: https://huggingface.co/datasets/surry-hills-druid/noagenda-transcripts.text1M<n<10M1 likes272 downloads2mo agoHugging Face25josemancharo /apptek_callcenter_dialogues_travel_hospitality_no_transcripts AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts) This is a filtered derivative of AppTek Call-Center Dialogues, prepared for a specific use case. Changes from the source dataset Restricted the dataset to the travel and hospitality domains. Removed the transcript field (text) entirely. Kept the original audio and the domain, gender, and accent metadata. Preserved the source dataset's test split. This dataset has transcripts removed and is… See the full description on the dataset page: https://huggingface.co/datasets/josemancharo/apptek_callcenter_dialogues_travel_hospitality_no_transcripts.audioaudio-classificationn<1K1 likes272 downloads2mo agoHugging Face26octava /indonesian-voice-transcription-1.4.85aaudio10K<n<100K0 likes259 downloads2y agoHugging Face27octava /indonesian-voice-transcription-1.3.9caudio10K<n<100K0 likes236 downloads2y agoHugging Face28WhissleAI /indicvoices_hi_tagged_transcripts Dataset Card for indicvoices_hi_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes228 downloads2y agoHugging Face29octava /indonesian-voice-transcription-1.4.9rcvaudio10K<n<100K0 likes213 downloads2y agoHugging Face30octava /indonesian-voice-transcription-1.4.9aaudio100K<n<1M0 likes203 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.