Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01japanese-asr /whisper_transcriptions.reazon_speech_all.wer_10.0.vectorized1M<n<10M0 likes85k downloads2y agoHugging Face02japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes48k downloads2y agoHugging Face03japanese-asr /whisper_transcriptions.mls.wer_10.0.vectorized1M<n<10M1 likes26k downloads2y agoHugging Face04japanese-asr /whisper_transcriptions.mls.wer_10.0audio1M<n<10M2 likes14k downloads2y agoHugging Face05japanese-asr /whisper_transcriptions.reazonspeech.allaudio10M<n<100M4 likes6.6k downloads2y agoHugging Face06japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0audio1M<n<10M3 likes3.4k downloads2y agoHugging Face07Whispering-GPT /linustechtips-transcript-audio Dataset Card for "linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.audioautomatic-speech-recognitionn<1K4 likes3k downloads4y agoHugging Face08Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes2.5k downloads4y agoHugging Face09nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B10 likes2.3k downloads2y agoHugging Face10lucazhou2000 /sciencemysterybench-transcriptsimagen<1K0 likes2.3k downloads22d agoHugging Face11ARTPARK-IISc /Vaani-transcription-partgatedThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages. This table represents the audio and transcription duration data for various languages. Language Angami Angika Ao Assamese Awadhi Bajjika Bearybashe Bengali Bhili Bhojpuri Bundeli Chakhesang Chakma Chhattisgarhi English Garhwali Garo Gondi Gujarati Halbi Haryanvi Hindi IduMishmi Kannada Kashmiri Karbi Khariboli Khortha Kokborok Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.audioautomatic-speech-recognition1M<n<10M20 likes2.2k downloads2d agoHugging Face12PaulR11 /petri-audit-transcripts-32q Petri Audit Transcripts (32 Quirks) — Qwen3 Baseline Corpus Baseline Petri audit transcripts used as the corpus for crux-eval construction and strategy-clustering analysis in the paper "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026). What's here ~8,000 audit transcripts produced by a baseline Qwen3-30B-A3B-Instruct-2507 auditor against Grok 4.1 Fast targets across 32 system-prompted model-organism quirks. One transcript per (quirk, seed, rollout).… See the full description on the dataset page: https://huggingface.co/datasets/PaulR11/petri-audit-transcripts-32q.0 likes1.5k downloads6mo agoHugging Face13ivrit-ai /audio-v2-transcriptsgated Overview This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It was released on May 18th, 2025. You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt. All files were transcribed using the process.py pipeline, performing: Frame-level VAD Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.audio-classification10K<n<100K1 likes1.4k downloads11mo agoHugging Face14kurry /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.tabulartext-generation10K<n<100K20 likes1.3k downloads1y agoHugging Face15nyu-dice-lab /wavepulse-radio-summarized-transcripts WavePulse Radio Summarized Transcripts Dataset Summary WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.texttext-generation100K<n<1M1 likes1.3k downloads2y agoHugging Face16comp-genomics-consortium /raw-transcriptome-unmapped-reads Dataset Card for CGSC Raw Transcriptome Unmapped Reads Dataset Summary This repository acts as the primary cold-storage for unmapped, raw sequencing outputs generated during the Q2 2026 Synthetic Bio-Arrays trials. The dataset comprises massive, uncompressed binary blobs that represent pre-alignment genomic data directly from the sequencing hardware. Because these files bypass standard alignment and compression algorithms (such as BAM/CRAM conversion) to preserve… See the full description on the dataset page: https://huggingface.co/datasets/comp-genomics-consortium/raw-transcriptome-unmapped-reads.tabular-classification100K<n<1M1 likes1.3k downloads4mo agoHugging Face17japanese-asr /whisper_transcriptions.mlsaudio10M<n<100M1 likes1.3k downloads2y agoHugging Face18ghanaopenai /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes1.2k downloads1mo agoHugging Face19japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0.vectorized1M<n<10M0 likes1.2k downloads2y agoHugging Face20jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face21Rogersurf /earnings-call-transcriptslanguage: en tags: finance earnings-calls transcripts nlp llm rag financial-analysis license: other pretty_name: Earnings Call Transcripts size_categories: - 10K<n<100K Earnings Call Transcripts Dataset A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages. Dataset Overview This dataset contains: Company earnings call transcripts Ticker symbols Earnings quarters Earnings years Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.tabular1K<n<10K1 likes919 downloads5mo agoHugging Face22My-Weird-Prompts /transcripts My Weird Prompts — Transcript Corpus Every published transcript from the My Weird Prompts podcast, shaped for textual analysis: narrowed metadata, the full transcript, the same transcript segmented into speaker turns, and per-episode text statistics. 5,601 episodes · 500,361 speaker turns. Rebuilt daily from the production database. Configs from datasets import load_dataset episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.tabulartext-generation100K<n<1M0 likes903 downloads20h agoHugging Face23Teejeigh /raw_friends_series_transcriptRaw transcript from friends tv series, chunk into ~1000 token lines. Text tagged by character. language: - en size_categories: - n<1K textn<1K2 likes887 downloads3y agoHugging Face24Bose345 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.tabulartext-generation10K<n<100K7 likes832 downloads10mo agoHugging Face25glopardo /sp500-earnings-transcripts S&P 500 Earnings Call Transcripts Dataset Description This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals. 📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025. Coverage Statistics Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.tabulartext-classification10K<n<100K7 likes786 downloads11mo agoHugging Face26ghananlpcommunity /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes771 downloads1mo agoHugging Face27openbank-uz /youtube_transcriptions Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google's Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: https://huggingface.co/datasets/openbank-uz/youtube_transcriptions.audioautomatic-speech-recognition100K<n<1M2 likes748 downloads7mo agoHugging Face28lytang /MeetingBank-transcriptThis dataset consists of transcripts from the MeetingBank dataset. Overview MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for… See the full description on the dataset page: https://huggingface.co/datasets/lytang/MeetingBank-transcript.textsummarization1K<n<10K19 likes738 downloads3y agoHugging Face29japanese-asr /whisper_transcriptions.reazon_speech_all.wer_10.0audio1M<n<10M0 likes714 downloads2y agoHugging Face30rmems /multi-agent-coordination-transcripts Multi Agent Coordination Transcripts Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.text1K<n<10K0 likes681 downloads20d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.