datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.earnings-call-transcriptslanguage:
en
tags:
finance
earnings-calls
transcripts
nlp
llm
rag
financial-analysis
license: other
pretty_name: Earnings Call Transcripts
size_categories:
- 10K<n<100K
Earnings Call Transcripts Dataset
A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages.
Dataset Overview
This dataset contains:
Company earnings call transcripts
Ticker symbols
Earnings quarters
Earnings years
Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.transcripts
My Weird Prompts — Transcript Corpus
Every published transcript from the My Weird Prompts
podcast, shaped for textual analysis: narrowed metadata, the full transcript, the
same transcript segmented into speaker turns, and per-episode text statistics.
5,690 episodes · 513,097 speaker turns. Rebuilt daily from the
production database.
Configs
from datasets import load_dataset
episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.sp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.sp500-earnings-transcripts
S&P 500 Earnings Call Transcripts
Dataset Description
This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals.
📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025.
Coverage Statistics
Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.kcc-farmer-query-transcriptstranscription-corpus
UN Transcription Corpus
Two splits of UN meeting audio paired with official verbatim records.
Splits
sessions — Whole meeting sessions (SC + GA plenary)
One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org.
Column
Description
symbol
UN document symbol, e.g. S/PV.9826
webtv_url
URL on UN Web TV
duration_ms
Session duration in milliseconds
num_speakers
Number of speaker turns in the verbatim record
audio_floor
Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.sejm-committee-transcripts
Polish Sejm committee transcripts — full API coverage (terms 9 and 10)
Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm
of the Republic of Poland, parsed into individually attributed speaker turns.
Scope
Committees: all standing committees with zapis PDFs in the Sejm API.
Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17).
Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.stk-earnings-transcriptssp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.LisTAya-transcripts
LisTAya transcripts: the test-set evaluations of the LisTAya study
This dataset holds the reference and the model output for every utterance of every test-set evaluation in the paper Does Regional Decoder Specialization Help Low-Resource ASR Based on the SLAM-ASR Framework? (ROCLING 2026). The trained models are described in the model card ERISLab/LisTAya and listed in the collection… See the full description on the dataset page: https://huggingface.co/datasets/ERISLab/LisTAya-transcripts.measuring_cot_monitorability_transcripts
Measuring Chain-of-Thought Monitorability Transcripts
This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness.
We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.salt-asr-data-transcriptionscollab-arena-v0-transcripts
Collaboration Arena v0 — transcripts & annotations
⚠️ This dataset is LIVE and GROWING until the experiment completes. Configs and counts update as Qwen cells and the 32B tier land; see the commit history for the update log.
ℹ️ E5 frontier model = claude-opus-4-8 (E1-E4 frontier = claude-fable-5). Fable safety-refused ~50% of E5 turns (stop_reason=refusal), concentrated on the honest cross-check seats, so E5's frontier arm was run on Opus instead; team and solo are BOTH Opus… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/collab-arena-v0-transcripts.fomc-meeting-transcripts
FOMC Meeting Transcripts (1976–2020)
Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board.
The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.untrivial-retrieval-transcripts
Untrivial Retrieval — Inspect Transcripts
Read the full Inspect transcripts.
The source code is here
DO NOT TRAIN OR CRAWL
Canary: UNTRIVIAL_RETRIEVAL_DO_NOT_TRAIN_OR_CRAWL_7B3A6E29C140
These transcripts are shared for human inspection, evaluation analysis, and
reproducibility. Do not use them for model training, fine-tuning, distillation,
or inclusion in training corpora. Do not crawl or bulk harvest this repository
or viewer for those purposes. Targeted… See the full description on the dataset page: https://huggingface.co/datasets/ConcurrentSquared/untrivial-retrieval-transcripts.vaani-all-eng-transcripts-concatsp500_earnings_transcripts
S&P 500 Earnings Transcripts Dataset
This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies.
Dataset Description
This collection includes:
Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.massive-yt-edu-transcriptions
Massive YouTube Educational Transcriptions
Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5.
Stats
Videos: 59,355
Characters: 1,539,022,925 (~384M tokens)
Audio hours: 35,890
Model: faster-whisper (CTranslate2) with distil-large-v3.5
Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime
Fields
Field
Description
video_id
YouTube video ID
title
Video title
text
Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.seamless-interact-canary-transcripts
Seamless Interact - Canary Transcripts
Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries.
Dataset Description
This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying… See the full description on the dataset page: https://huggingface.co/datasets/hirotakahiraki/seamless-interact-canary-transcripts.swe_chat_transcriptsKisan_Call_Centre_Transcriptshome-telecom-callcenter-transcripts-viewer
Home/Telecom Call Center Transcript Viewer V2
This public dataset is a flattened viewer copy of one source archive from AIxBlock/92k-real-world-call-center-scripts-english:
home_ervice_inbound&telecom _outbound.zip
It was converted so Hugging Face Data Studio can display the contents as regular Parquet tables.
Splits
default/train: one row per transcript, all 3,239 transcripts from the source zip.
turns/turns: one row per inferred timestamp/sentence turn, linked by… See the full description on the dataset page: https://huggingface.co/datasets/yevgeniy03/home-telecom-callcenter-transcripts-viewer.voynich-transcription-mismatch
Voynich Transcription Mismatch Index
Dataset Summary
This dataset provides a line-by-line comparison across five different transcription sources of the Voynich Manuscript. It tracks agreements and disagreements between transcribers, enabling research on transcription uncertainty and consensus.
The primary comparison is between EVA-based transcriptions (ZL and IT), with additional tracking of Currier, FSG, and v101 transcription systems.
Key Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Ched-ai/voynich-transcription-mismatch.earnings_call_transcript_autograder
Earnings Call LLM Insights
📚 Read the Full Story: For a deep dive into the methodology, the wildest moments we found, and key takeaways, check out the blog post:KnowTrend.ai: Auto-Grading Ten Years of Earnings Calls for Prescience and Delusion
This dataset contains LLM-generated analysis of ~70,000+ earnings call transcripts.
The analysis was performed using Kimi k2-0905-preview, focusing on extracting specific insights, prescient analyst questions, and management missteps.… See the full description on the dataset page: https://huggingface.co/datasets/knowtrendllc/earnings_call_transcript_autograder.agenttime-transcripts
AgentTime transcripts
Agent transcripts from the experiments in "AgentTime: Can Agents Estimate and Control Their Own Runtime?"
(paper, website, code). Each row is one run,
forecast or follow-up question, with its native Claude Code, Codex, Claude Agent SDK or OpenRouter transcript. The
transcripts are scrubbed of credentials and personal or infrastructure identifiers. Task text from four benchmarks is redacted.
There are 8,911 rows in 12 configs.
BENCHMARK DATA SHOULD NEVER… See the full description on the dataset page: https://huggingface.co/datasets/mofengenden/agenttime-transcripts.podcast-transcripts
Podcast Transcripts Dataset
This dataset contains transcripts from Bitcoin and cryptocurrency podcasts,
processed by the belief-engines ETL pipeline.
Files
transcripts.parquet - Full episode transcripts with speaker diarization
transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings
Schema
transcripts.parquet
episode_id: Unique episode identifier
podcast_slug: Podcast name slug
episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.filled_stacks_transcriptions
Dataset Card for "filled_stacks_transcriptions"
More Information needed
meld-transcript-final
