Team Ai
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B10 likes1.7k downloads2y agoHugging Face02kurry /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.tabulartext-generation10K<n<100K20 likes1.3k downloads1y agoHugging Face03My-Weird-Prompts /transcripts My Weird Prompts — Transcript Corpus Every published transcript from the My Weird Prompts podcast, shaped for textual analysis: narrowed metadata, the full transcript, the same transcript segmented into speaker turns, and per-episode text statistics. 5,690 episodes · 513,097 speaker turns. Rebuilt daily from the production database. Configs from datasets import load_dataset episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.tabulartext-generation100K<n<1M0 likes934 downloads12h agoHugging Face04Bose345 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.tabulartext-generation10K<n<100K7 likes823 downloads10mo agoHugging Face05PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes236 downloads23d agoHugging Face06churchill1254 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/churchill1254/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes157 downloads6mo agoHugging Face07ameek /measuring_cot_monitorability_transcripts Measuring Chain-of-Thought Monitorability Transcripts This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness. We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.tabularquestion-answering100K<n<1M1 likes125 downloads11mo agoHugging Face08lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes122 downloads1mo agoHugging Face09lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes113 downloads1mo agoHugging Face10sahar-millis-runi /old-games-transcript 90s Games Transcript A small English-language corpus of narrative text from classic PC games. This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading. Dataset Config Records Content caesar3 20 Mission briefings + victory messages diablo2_lod 7 Cinematic narration and dialogue warcraft2 53 Human + Orc… See the full description on the dataset page: https://huggingface.co/datasets/sahar-millis-runi/old-games-transcript.tabulartext-generationn<1K1 likes84 downloads13d agoHugging Face11AltaySec /altayduel-transcripts 🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2) Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir. 📌 TL;DR 2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu. 439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.tabulartext-classification1K<n<10K1 likes75 downloads2mo agoHugging Face12brishen /fomc-meeting-transcripts FOMC Meeting Transcripts (1976–2020) Full-text transcripts of 373 Federal Open Market Committee (FOMC) meetings, from March 1976 through December 2020, converted from the official PDF transcripts published by the Federal Reserve Board. The FOMC is the body of the U.S. Federal Reserve System that sets monetary policy (the federal funds rate target, balance-sheet policy, etc.). Verbatim meeting transcripts are released to the public with a roughly five-year lag, which is why… See the full description on the dataset page: https://huggingface.co/datasets/brishen/fomc-meeting-transcripts.tabulartext-generationn<1K0 likes67 downloads1mo agoHugging Face13idleengine /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/sp500_earnings_transcripts.tabulartext-generation10K<n<100K0 likes58 downloads2mo agoHugging Face14thepowerfuldeez /massive-yt-edu-transcriptions Massive YouTube Educational Transcriptions Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5. Stats Videos: 59,355 Characters: 1,539,022,925 (~384M tokens) Audio hours: 35,890 Model: faster-whisper (CTranslate2) with distil-large-v3.5 Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime Fields Field Description video_id YouTube video ID title Video title text Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.tabularautomatic-speech-recognition10K<n<100K3 likes51 downloads5mo agoHugging Face15Infektyd /council-transcripts Council Multi-Agent Deliberation Transcripts Real deliberation transcripts from Council, a multi-agent orchestration skill for OpenClaw. What This Is Council routes tasks to specialized Grok persona agents — Workhorse (deep technical reasoning), Creative (novel ideas, chaos energy), and Speed (fast iteration) — then synthesizes their outputs into a final verdict via a Conductor. These transcripts capture the full deliberation process: prompts, per-persona responses, and… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/council-transcripts.tabulartext-generationn<1K0 likes41 downloads6mo agoHugging Face16mofengenden /agenttime-transcriptsgated AgentTime transcripts Agent transcripts from the experiments in "AgentTime: Can Agents Estimate and Control Their Own Runtime?" (paper, website, code). Each row is one run, forecast or follow-up question, with its native Claude Code, Codex, Claude Agent SDK or OpenRouter transcript. The transcripts are scrubbed of credentials and personal or infrastructure identifiers. Task text from four benchmarks is redacted. There are 8,911 rows in 12 configs. BENCHMARK DATA SHOULD NEVER… See the full description on the dataset page: https://huggingface.co/datasets/mofengenden/agenttime-transcripts.tabulartext-generation1K<n<10K2 likes36 downloads3d agoHugging Face17rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes34 downloads8mo agoHugging Face18kaushik-systalyze /customer-transcript-short-control Customer Transcript Short Control Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-short-control.tabulartext-generation1K<n<10K0 likes28 downloads4mo agoHugging Face19kaushik-systalyze /customer-transcript-analytics Customer Transcript Analytics Curated customer-support and meeting transcripts mapped to a single fixed "analyze this transcript → compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows range from short support chats… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-analytics.tabulartext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face20kaushik-systalyze /customer-transcript-holdout-eval Customer Transcript Holdout Eval Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-holdout-eval.tabulartext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face21kaushik-systalyze /customer-transcript-long-dialog Customer Transcript Long Dialogue Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-long-dialog.tabulartext-generation1K<n<10K0 likes21 downloads4mo agoHugging Face22paulalesius /terence-mckenna-transcripts Terence McKenna Transcripts Catalog of machine-transcribed talks and interviews, mostly by Terence McKenna, built from 172 videos. Two tables: talks — one row per video; full original transcript (including [SPEAKER_XX] diarization tags) turns — one row per non-empty line; speaker tags stripped from text Speaker labels are not consistent across files (no cross-video voice fingerprinting), so turns has no speaker column. Load from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/paulalesius/terence-mckenna-transcripts.tabulartext-generation100K<n<1M0 likes20 downloads1mo agoHugging Face23kaushik-systalyze /customer-transcript-source Customer Transcript Source Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token accounting… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-source.tabulartext-generation1K<n<10K0 likes18 downloads4mo agoHugging Face24PiotrSty /yodas3-transcripts-filipino YODAS Filipino transcripts (yodas3_fil) - spoken-language captions Filipino caption tracks from YODAS v3, one complete transcript per record. This text-only snapshot covers every Filipino metadata shard at the pinned source revision. English translations and word-level duplicate annotations are excluded from the text. Source records: 6,790 across 16 shards Retained records: 6,565 (96.7%) Measured cl100k_base proxy tokens: 21,565,408 Characters: 63,775,326 Before exact… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/yodas3-transcripts-filipino.tabulartext-generation1K<n<10K0 likes15 downloads3d agoHugging Face25tonychenxyz /frontier-ai-podcast-transcripts Frontier AI Researcher Podcast Transcripts Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode. Contents 327 episodes 49,286 merged dialogue turns 5,595,982 English tokens using the o200k_base tokenizer 3,942,026 tokens in guest turns Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.tabulartext-generationn<1K0 likes14 downloads2mo agoHugging Face26build-small-hackathon /velvet-rope-playtest-transcripts Velvet Rope Playtest Transcripts Cleaned playtest transcripts for Velvet Rope, a Build Small Hackathon Gradio game where players talk past whimsical AI gatekeepers by reading moods and discovering each character's soft spot. This dataset is published for the hackathon's sharing-is-caring badge. It contains 341 turn-level rows from 96 local playtest session files. Files data/playtest_transcripts.csv - table-friendly version. data/playtest_transcripts.jsonl - one… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/velvet-rope-playtest-transcripts.tabulartext-generationn<1K0 likes13 downloads4mo agoHugging Face27hudsongouge /podcast-transcripts-cleaned-phase2gated Podcast Transcripts Cleaned (Phase 2) Updated 2026-07-18T23-19-55Z UTC. Configs Config Rows Description episodes 4,304 Full-episode raw ASR → cleaned transcript (with episode_id / show_id) chunks 5,646 Per-chunk raw → cleaned pairs traces 287,271 Full LM traces (cleaner + evaluator): prompts, reasoning/CoT, tool outputs, pass/fail Cleaned pairs (episodes / chunks) Columns include episode_id, show_id, title, url, instruction… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/podcast-transcripts-cleaned-phase2.tabulartext-generation100K<n<1M1 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.