Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B10 likes2.2k downloads2y agoHugging Face02diarizers-community /ami_ihm_with_transcriptsaudion<1K0 likes292 downloads2y agoHugging Face03surry-hills-druid /noagenda-transcripts noagenda transcripts This is the dataset for transcripts of the noagendashow.net podcast. The transcripts are in the data folder. It also contains the source code for generating the transcripts, and the source code for the noagenda-transcripts.net website which searches the transcripts and plays the audio clips for each search result and takes you to the location in the transcript. The code in the repo consists of 2 main parts: A Go CLI for transcribing the audio, creating the… See the full description on the dataset page: https://huggingface.co/datasets/surry-hills-druid/noagenda-transcripts.text1M<n<10M1 likes272 downloads2mo agoHugging Face04josemancharo /apptek_callcenter_dialogues_travel_hospitality_no_transcripts AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts) This is a filtered derivative of AppTek Call-Center Dialogues, prepared for a specific use case. Changes from the source dataset Restricted the dataset to the travel and hospitality domains. Removed the transcript field (text) entirely. Kept the original audio and the domain, gender, and accent metadata. Preserved the source dataset's test split. This dataset has transcripts removed and is… See the full description on the dataset page: https://huggingface.co/datasets/josemancharo/apptek_callcenter_dialogues_travel_hospitality_no_transcripts.audioaudio-classificationn<1K1 likes272 downloads2mo agoHugging Face05WhissleAI /indicvoices_hi_tagged_transcripts Dataset Card for indicvoices_hi_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes228 downloads2y agoHugging Face06WhissleAI /indicvoices_pa_tagged_transcripts Dataset Card for indicvoices_pa_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes188 downloads2y agoHugging Face07prince-canuma /accentsDB-with-transcripts Dataset Card for "accentsDB-with-transcripts" More Information needed audio10K<n<100K0 likes42 downloads3y agoHugging Face08rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes33 downloads8mo agoHugging Face09quangdung /gigaspeech2-vi-missing-transcripts GigaSpeech2 Vietnamese WAVs Missing Transcripts This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs are absent from train_refined.tsv. Contents 201,295 WAV files without a matching transcript 41 uncompressed TAR shards in shards/ 36 GB of audio (approximately) processing_manifest.jsonl: per-source-archive counts summary.json: aggregate counts SHA256SUMS: checksums for all TAR shards Each shard preserves the original relative layout:… See the full description on the dataset page: https://huggingface.co/datasets/quangdung/gigaspeech2-vi-missing-transcripts.audioautomatic-speech-recognition0 likes26 downloads2mo agoHugging Face10WhissleAI /indicvoices_bn_tagged_transcripts Dataset Card for indicvoices_bn_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_bn_tagged_transcripts.audioautomatic-speech-recognitionn<1K0 likes21 downloads2y agoHugging Face11WhissleAI /indicvoices_mr_tagged_transcripts Dataset Card for indicvoices_mr_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_mr_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes12 downloads2y agoHugging Face12komalgupta23 /medical-audio-transcriptsaudio1K<n<10K0 likes12 downloads4mo agoHugging Face13ananyakarn /Androids-Corpus-with-Transcriptsaudio0 likes9 downloads11mo agoHugging Face14DarAudy /SA_transcripts license: apache-2.0 ---This is a test dataset audioaudio-classificationn<1K0 likes6 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.