Team Ai
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes441 downloads7mo agoHugging Face02ERISLab /LisTAya-transcripts LisTAya transcripts: the test-set evaluations of the LisTAya study This dataset holds the reference and the model output for every utterance of every test-set evaluation in the paper Does Regional Decoder Specialization Help Low-Resource ASR Based on the SLAM-ASR Framework? (ROCLING 2026). The trained models are described in the model card ERISLab/LisTAya and listed in the collection… See the full description on the dataset page: https://huggingface.co/datasets/ERISLab/LisTAya-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes157 downloads5d agoHugging Face03thepowerfuldeez /massive-yt-edu-transcriptions Massive YouTube Educational Transcriptions Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5. Stats Videos: 59,355 Characters: 1,539,022,925 (~384M tokens) Audio hours: 35,890 Model: faster-whisper (CTranslate2) with distil-large-v3.5 Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime Fields Field Description video_id YouTube video ID title Video title text Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.tabularautomatic-speech-recognition10K<n<100K3 likes51 downloads5mo agoHugging Face04hirotakahiraki /seamless-interact-canary-transcripts Seamless Interact - Canary Transcripts Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries. Dataset Description This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying… See the full description on the dataset page: https://huggingface.co/datasets/hirotakahiraki/seamless-interact-canary-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes51 downloads7mo agoHugging Face05rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes34 downloads8mo agoHugging Face06hudsongouge /Podcast-Transcripts-Dedupedgated Podcast Transcripts (Deduped) Phase 1 deduplicated subset of hudsongouge/Podcast-Transcripts-Raw. Updated 2026-07-19T06-28-19Z UTC. Dedup summary Metric Value Raw input rows 102,374 Keepers 99,035 Discarded 3,339 Exact discarded 882 MinHash discarded 2457 Jaccard threshold 0.88 Two-pass Phase 1: Exact match on canonical audio source key (YouTube ID / Omny clip / audio filename) MinHash + LSH on normalized transcript text (Jaccard ≈… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Deduped.tabularautomatic-speech-recognition100K<n<1M1 likes9 downloads3mo agoHugging Face07hudsongouge /Podcast-Transcripts-Rawgated Podcast Transcripts Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio). This dataset will be gated. Only people who are part of our team may access. Splits Config Rows Description shows 124 Channels / podcast feeds (name, description, hosts, links) episodes 102,374 Episode/video metadata (title, description, guests, tags, dates) transcripts 102,374 ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.tabularautomatic-speech-recognition100K<n<1M1 likes8 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.