Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes34k downloads2y agoHugging Face02japanese-asr /whisper_transcriptions.reazonspeech.allaudio10M<n<100M4 likes4.5k downloads2y agoHugging Face03ARTPARK-IISc /Vaani-transcription-partgated 📢 Update (08 October 2026): The dataset has been updated with the following changes: Transcription corrections: Refined part of the transriptions. Language normalization: Language names have been normalized . Increased duration: The total transcribed speech duration has increased from 2048 to 2122 hours. This dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2122.20 hours, covering 86 languages. This table represents the audio… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.audioautomatic-speech-recognition1M<n<10M20 likes2.8k downloads8h agoHugging Face04japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0audio1M<n<10M3 likes2.6k downloads2y agoHugging Face05Whispering-GPT /linustechtips-transcript-audio Dataset Card for "linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.audioautomatic-speech-recognitionn<1K4 likes2.5k downloads4y agoHugging Face06nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B10 likes1.7k downloads2y agoHugging Face07Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes1.7k downloads4y agoHugging Face08kurry /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.tabulartext-generation10K<n<100K20 likes1.3k downloads1y agoHugging Face09ghanaopenai /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes1.2k downloads1mo agoHugging Face10Rogersurf /earnings-call-transcriptslanguage: en tags: finance earnings-calls transcripts nlp llm rag financial-analysis license: other pretty_name: Earnings Call Transcripts size_categories: - 10K<n<100K Earnings Call Transcripts Dataset A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages. Dataset Overview This dataset contains: Company earnings call transcripts Ticker symbols Earnings quarters Earnings years Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.tabular1K<n<10K1 likes951 downloads5mo agoHugging Face11japanese-asr /whisper_transcriptions.mlsaudio10M<n<100M1 likes934 downloads2y agoHugging Face12My-Weird-Prompts /transcripts My Weird Prompts — Transcript Corpus Every published transcript from the My Weird Prompts podcast, shaped for textual analysis: narrowed metadata, the full transcript, the same transcript segmented into speaker turns, and per-episode text statistics. 5,690 episodes · 513,097 speaker turns. Rebuilt daily from the production database. Configs from datasets import load_dataset episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.tabulartext-generation100K<n<1M0 likes934 downloads3h agoHugging Face13nyu-dice-lab /wavepulse-radio-summarized-transcripts WavePulse Radio Summarized Transcripts Dataset Summary WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.texttext-generation100K<n<1M1 likes872 downloads2y agoHugging Face14Bose345 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.tabulartext-generation10K<n<100K7 likes823 downloads10mo agoHugging Face15glopardo /sp500-earnings-transcripts S&P 500 Earnings Call Transcripts Dataset Description This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals. 📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025. Coverage Statistics Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.tabulartext-classification10K<n<100K7 likes781 downloads11mo agoHugging Face16openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes759 downloads7mo agoHugging Face17japanese-asr /whisper_transcriptions.reazonspeech.largeaudio1M<n<10M0 likes549 downloads3y agoHugging Face18AnmolNimmala0 /kcc-farmer-query-transcriptstabular10M<n<100M0 likes544 downloads8d agoHugging Face19ghananlpcommunity /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes538 downloads1mo agoHugging Face20united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes441 downloads7mo agoHugging Face21japanese-asr /whisper_transcriptions.reazonspeech.mediumaudio100K<n<1M0 likes431 downloads3y agoHugging Face22yuriyvnv /synthetic_transcript_pt Portuguese Speech Dataset with Multiple Training Configurations A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms. 🎯 Dataset Configurations Overview This dataset provides three carefully curated subsets to enable comprehensive speech recognition research: Configuration Training Data Validation Test Total Samples Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.audioautomatic-speech-recognition100K<n<1M0 likes410 downloads5mo agoHugging Face23modulate /entity-transcription-benchmark Entity Transcription Benchmark Measures whether a speech recognition system transcribes named entities correctly — as distinct from word error rate. WER weights every token equally. The tokens that matter for redaction, lookup, routing and search are proper nouns, and they are a small fraction of any transcript. A system can improve WER while getting worse at exactly the words a downstream consumer needs, and nothing in the standard evaluation will show it. 2,151 clips, 6.0… See the full description on the dataset page: https://huggingface.co/datasets/modulate/entity-transcription-benchmark.audioautomatic-speech-recognition1K<n<10K4 likes373 downloads29d agoHugging Face24japanese-asr /whisper_transcriptions.reazonspeech.large.wer_10.0audio1M<n<10M0 likes357 downloads3y agoHugging Face25octava /indonesian-voice-transcription-1.4.9a-raudio10K<n<100K0 likes323 downloads2y agoHugging Face26diarizers-community /ami_ihm_with_transcriptsaudion<1K0 likes274 downloads2y agoHugging Face27metr-evals /malt-transcripts-publicgated MALT: Manually-Reviewed Agentic Labeled Transcripts MALT-public is our collection of agent transcripts. Our public variant only includes data on non-internal tasks, which includes 30 task families and 169 tasks, across ~19 different models (some might be different releases of the same model, from different providers, or internal naming changes). Here's a summary table: has_chain_of_thought labels model manually_reviewed run_source count False bypass_constraints… See the full description on the dataset page: https://huggingface.co/datasets/metr-evals/malt-transcripts-public.text10K<n<100K7 likes261 downloads7mo agoHugging Face28octava /indonesian-voice-transcription-1.4.9a.2audio10K<n<100K0 likes255 downloads2y agoHugging Face29Whispering-GPT /yannick-kilcher-transcript-audio Dataset Card for "yannic-kilcher-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher. Data Fields The dataset is composed by: id: Id of the youtube video. channel:… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannick-kilcher-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes247 downloads4y agoHugging Face30PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes236 downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.