Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B10 likes1.7k downloads2y agoHugging Face02kurry /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.tabulartext-generation10K<n<100K20 likes1.3k downloads1y agoHugging Face03Rogersurf /earnings-call-transcriptslanguage: en tags: finance earnings-calls transcripts nlp llm rag financial-analysis license: other pretty_name: Earnings Call Transcripts size_categories: - 10K<n<100K Earnings Call Transcripts Dataset A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages. Dataset Overview This dataset contains: Company earnings call transcripts Ticker symbols Earnings quarters Earnings years Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.tabular1K<n<10K1 likes951 downloads5mo agoHugging Face04My-Weird-Prompts /transcripts My Weird Prompts — Transcript Corpus Every published transcript from the My Weird Prompts podcast, shaped for textual analysis: narrowed metadata, the full transcript, the same transcript segmented into speaker turns, and per-episode text statistics. 5,690 episodes · 513,097 speaker turns. Rebuilt daily from the production database. Configs from datasets import load_dataset episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.tabulartext-generation100K<n<1M0 likes934 downloads8h agoHugging Face05Bose345 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.tabulartext-generation10K<n<100K7 likes823 downloads10mo agoHugging Face06glopardo /sp500-earnings-transcripts S&P 500 Earnings Call Transcripts Dataset Description This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals. 📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025. Coverage Statistics Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.tabulartext-classification10K<n<100K7 likes781 downloads11mo agoHugging Face07AnmolNimmala0 /kcc-farmer-query-transcriptstabular10M<n<100M0 likes544 downloads8d agoHugging Face08united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes441 downloads7mo agoHugging Face09PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes236 downloads22d agoHugging Face10deerfieldgreen /stk-earnings-transcriptstabular1K<n<10K0 likes180 downloads2y agoHugging Face11churchill1254 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes157 downloads6mo agoHugging Face12ERISLab /LisTAya-transcripts LisTAya transcripts: the test-set evaluations of the LisTAya study This dataset holds the reference and the model output for every utterance of every test-set evaluation in the paper Does Regional Decoder Specialization Help Low-Resource ASR Based on the SLAM-ASR Framework? (ROCLING 2026). The trained models are described in the model card ERISLab/LisTAya and listed in the collection… See the full description on the dataset page: https://huggingface.co/datasets/ERISLab/LisTAya-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes157 downloads5d agoHugging Face13ameek /measuring_cot_monitorability_transcripts Measuring Chain-of-Thought Monitorability Transcripts This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness. We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.tabularquestion-answering100K<n<1M1 likes125 downloads11mo agoHugging Face14jq /salt-asr-data-transcriptionstabular10K<n<100K0 likes120 downloads2y agoHugging Face15siddharthmb /collab-arena-v0-transcripts Collaboration Arena v0 — transcripts & annotations ⚠️ This dataset is LIVE and GROWING until the experiment completes. Configs and counts update as Qwen cells and the 32B tier land; see the commit history for the update log. ℹ️ E5 frontier model = claude-opus-4-8 (E1-E4 frontier = claude-fable-5). Fable safety-refused ~50% of E5 turns (stop_reason=refusal), concentrated on the honest cross-check seats, so E5's frontier arm was run on Opus instead; team and solo are BOTH Opus… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/collab-arena-v0-transcripts.tabularother1K<n<10K0 likes115 downloads3mo agoHugging Face16brishen /fomc-meeting-transcripts FOMC Meeting Transcripts (1976–2020) Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board. The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.tabulartext-generationn<1K0 likes67 downloads1mo agoHugging Face17ConcurrentSquared /untrivial-retrieval-transcripts Untrivial Retrieval — Inspect Transcripts Read the full Inspect transcripts. The source code is here DO NOT TRAIN OR CRAWL Canary: UNTRIVIAL_RETRIEVAL_DO_NOT_TRAIN_OR_CRAWL_7B3A6E29C140 These transcripts are shared for human inspection, evaluation analysis, and reproducibility. Do not use them for model training, fine-tuning, distillation, or inclusion in training corpora. Do not crawl or bulk harvest this repository or viewer for those purposes. Targeted… See the full description on the dataset page: https://huggingface.co/datasets/ConcurrentSquared/untrivial-retrieval-transcripts.tabularn<1K0 likes65 downloads7d agoHugging Face18dianavdavidson /vaani-all-eng-transcripts-concattabular10K<n<100K0 likes61 downloads4mo agoHugging Face19idleengine /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes58 downloads2mo agoHugging Face20thepowerfuldeez /massive-yt-edu-transcriptions Massive YouTube Educational Transcriptions Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5. Stats Videos: 59,355 Characters: 1,539,022,925 (~384M tokens) Audio hours: 35,890 Model: faster-whisper (CTranslate2) with distil-large-v3.5 Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime Fields Field Description video_id YouTube video ID title Video title text Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.tabularautomatic-speech-recognition10K<n<100K3 likes51 downloads5mo agoHugging Face21hirotakahiraki /seamless-interact-canary-transcripts Seamless Interact - Canary Transcripts Speech transcription dataset generated by running NVIDIA Canary-Qwen2.5B ASR model on the Seamless Interact conversational speech dataset, segmented by original corpus boundaries. Dataset Description This dataset contains 2,781,985 transcribed speech segments from the Seamless Interact corpus. Each segment includes the transcribed text, timing information (offset and duration within the source audio), and metadata identifying… See the full description on the dataset page: https://huggingface.co/datasets/hirotakahiraki/seamless-interact-canary-transcripts.tabularautomatic-speech-recognition1M<n<10M0 likes51 downloads7mo agoHugging Face22rguan72 /swe_chat_transcriptstabularn<1K0 likes51 downloads5mo agoHugging Face23Omegaindebt /Kisan_Call_Centre_Transcriptstabular1M<n<10M0 likes50 downloads1y agoHugging Face24yevgeniy03 /home-telecom-callcenter-transcripts-viewer Home/Telecom Call Center Transcript Viewer V2 This public dataset is a flattened viewer copy of one source archive from AIxBlock/92k-real-world-call-center-scripts-english: home_ervice_inbound&telecom _outbound.zip It was converted so Hugging Face Data Studio can display the contents as regular Parquet tables. Splits default/train: one row per transcript, all 3,239 transcripts from the source zip. turns/turns: one row per inferred timestamp/sentence turn, linked by… See the full description on the dataset page: https://huggingface.co/datasets/yevgeniy03/home-telecom-callcenter-transcripts-viewer.tabular100K<n<1M0 likes50 downloads6mo agoHugging Face25Ched-ai /voynich-transcription-mismatch Voynich Transcription Mismatch Index Dataset Summary This dataset provides a line-by-line comparison across five different transcription sources of the Voynich Manuscript. It tracks agreements and disagreements between transcribers, enabling research on transcription uncertainty and consensus. The primary comparison is between EVA-based transcriptions (ZL and IT), with additional tracking of Currier, FSG, and v101 transcription systems. Key Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Ched-ai/voynich-transcription-mismatch.tabularother1K<n<10K0 likes45 downloads4mo agoHugging Face26knowtrendllc /earnings_call_transcript_autograder Earnings Call LLM Insights 📚 Read the Full Story: For a deep dive into the methodology, the wildest moments we found, and key takeaways, check out the blog post:KnowTrend.ai: Auto-Grading Ten Years of Earnings Calls for Prescience and Delusion This dataset contains LLM-generated analysis of ~70,000+ earnings call transcripts. The analysis was performed using Kimi k2-0905-preview, focusing on extracting specific insights, prescient analyst questions, and management missteps.… See the full description on the dataset page: https://huggingface.co/datasets/knowtrendllc/earnings_call_transcript_autograder.tabular10K<n<100K0 likes37 downloads10mo agoHugging Face27mofengenden /agenttime-transcriptsgated AgentTime transcripts Agent transcripts from the experiments in "AgentTime: Can Agents Estimate and Control Their Own Runtime?" (paper, website, code). Each row is one run, forecast or follow-up question, with its native Claude Code, Codex, Claude Agent SDK or OpenRouter transcript. The transcripts are scrubbed of credentials and personal or infrastructure identifiers. Task text from four benchmarks is redacted. There are 8,911 rows in 12 configs. BENCHMARK DATA SHOULD NEVER… See the full description on the dataset page: https://huggingface.co/datasets/mofengenden/agenttime-transcripts.tabulartext-generation1K<n<10K2 likes36 downloads2d agoHugging Face28rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes34 downloads8mo agoHugging Face29bejaeger /filled_stacks_transcriptions Dataset Card for "filled_stacks_transcriptions" More Information needed tabular10K<n<100K0 likes32 downloads4y agoHugging Face30TAESOO98 /meld-transcript-finaltabular10K<n<100K1 likes32 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.