datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.ami_ihm_with_transcriptsnoagenda-transcripts
noagenda transcripts
This is the dataset for transcripts of the noagendashow.net podcast. The transcripts are in the data folder.
It also contains the source code for generating the transcripts, and the source code for the noagenda-transcripts.net website which searches the transcripts and plays the audio clips for each search result and takes you to the location in the transcript.
The code in the repo consists of 2 main parts:
A Go CLI for transcribing the audio, creating the… See the full description on the dataset page: https://huggingface.co/datasets/surry-hills-druid/noagenda-transcripts.apptek_callcenter_dialogues_travel_hospitality_no_transcripts
AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts)
This is a filtered derivative of AppTek Call-Center Dialogues, prepared for a specific use case.
Changes from the source dataset
Restricted the dataset to the travel and hospitality domains.
Removed the transcript field (text) entirely.
Kept the original audio and the domain, gender, and accent metadata.
Preserved the source dataset's test split.
This dataset has transcripts removed and is… See the full description on the dataset page: https://huggingface.co/datasets/josemancharo/apptek_callcenter_dialogues_travel_hospitality_no_transcripts.indicvoices_hi_tagged_transcripts
Dataset Card for indicvoices_hi_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.indicvoices_pa_tagged_transcripts
Dataset Card for indicvoices_pa_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.accentsDB-with-transcripts
Dataset Card for "accentsDB-with-transcripts"
More Information needed
podcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.gigaspeech2-vi-missing-transcripts
GigaSpeech2 Vietnamese WAVs Missing Transcripts
This dataset contains the Vietnamese GigaSpeech2 training WAV files whose IDs
are absent from train_refined.tsv.
Contents
201,295 WAV files without a matching transcript
41 uncompressed TAR shards in shards/
36 GB of audio (approximately)
processing_manifest.jsonl: per-source-archive counts
summary.json: aggregate counts
SHA256SUMS: checksums for all TAR shards
Each shard preserves the original relative layout:… See the full description on the dataset page: https://huggingface.co/datasets/quangdung/gigaspeech2-vi-missing-transcripts.indicvoices_bn_tagged_transcripts
Dataset Card for indicvoices_bn_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_bn_tagged_transcripts.indicvoices_mr_tagged_transcripts
Dataset Card for indicvoices_mr_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_mr_tagged_transcripts.medical-audio-transcriptsAndroids-Corpus-with-TranscriptsSA_transcripts
license: apache-2.0
---This is a test dataset
