Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucazhou2000 /sciencemysterybench-transcriptsimagen<1K0 likes2.3k downloads23d agoHugging Face02nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B10 likes2.2k downloads2y agoHugging Face03ivrit-ai /audio-v2-transcriptsgated Overview This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It was released on May 18th, 2025. You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt. All files were transcribed using the process.py pipeline, performing: Frame-level VAD Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.audio-classification10K<n<100K1 likes1.5k downloads11mo agoHugging Face04PaulR11 /petri-audit-transcripts-32q Petri Audit Transcripts (32 Quirks) — Qwen3 Baseline Corpus Baseline Petri audit transcripts used as the corpus for crux-eval construction and strategy-clustering analysis in the paper "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026). What's here ~8,000 audit transcripts produced by a baseline Qwen3-30B-A3B-Instruct-2507 auditor against Grok 4.1 Fast targets across 32 system-prompted model-organism quirks. One transcript per (quirk, seed, rollout).… See the full description on the dataset page: https://huggingface.co/datasets/PaulR11/petri-audit-transcripts-32q.0 likes1.5k downloads6mo agoHugging Face05kurry /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/kurry/sp500_earnings_transcripts.tabulartext-generation10K<n<100K20 likes1.4k downloads1y agoHugging Face06nyu-dice-lab /wavepulse-radio-summarized-transcripts WavePulse Radio Summarized Transcripts Dataset Summary WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.texttext-generation100K<n<1M1 likes1.3k downloads2y agoHugging Face07Rogersurf /earnings-call-transcriptslanguage: en tags: finance earnings-calls transcripts nlp llm rag financial-analysis license: other pretty_name: Earnings Call Transcripts size_categories: - 10K<n<100K Earnings Call Transcripts Dataset A cleaned financial NLP dataset containing earnings call transcripts collected from publicly available earnings call pages. Dataset Overview This dataset contains: Company earnings call transcripts Ticker symbols Earnings quarters Earnings years Call dates… See the full description on the dataset page: https://huggingface.co/datasets/Rogersurf/earnings-call-transcripts.tabular1K<n<10K1 likes923 downloads5mo agoHugging Face08My-Weird-Prompts /transcripts My Weird Prompts — Transcript Corpus Every published transcript from the My Weird Prompts podcast, shaped for textual analysis: narrowed metadata, the full transcript, the same transcript segmented into speaker turns, and per-episode text statistics. 5,608 episodes · 501,215 speaker turns. Rebuilt daily from the production database. Configs from datasets import load_dataset episodes = load_dataset("My-Weird-Prompts/transcripts", "episodes", split="train") # one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/transcripts.tabulartext-generation100K<n<1M0 likes922 downloads12h agoHugging Face09Bose345 /sp500_earnings_transcripts S&P 500 Earnings Transcripts Dataset This comprehensive dataset contains earnings call transcripts for S&P 500 companies and US large-caps, spanning from 2005 to 2025. Earnings calls provide valuable insights into company performance, strategic initiatives, and management perspectives that are essential for financial analysis, natural language processing research, and market sentiment studies. Dataset Description This collection includes: Complete transcripts: Full… See the full description on the dataset page: https://huggingface.co/datasets/Bose345/sp500_earnings_transcripts.tabulartext-generation10K<n<100K7 likes824 downloads10mo agoHugging Face10glopardo /sp500-earnings-transcripts S&P 500 Earnings Call Transcripts Dataset Description This dataset provides earnings call transcripts for S&P 500 companies, primarily covering 2014-2024, along with quarterly financial metrics and company fundamentals. 📄 Paper: This dataset was prepared for and used in Ca'Zorzi, Manu, Lopardo. Verba Volant, Transcripta Manent: What Corporate Earnings Calls Reveal About the AI Stock Rally. No. 3093. European Central Bank, 2025. Coverage Statistics Time… See the full description on the dataset page: https://huggingface.co/datasets/glopardo/sp500-earnings-transcripts.tabulartext-classification10K<n<100K7 likes808 downloads11mo agoHugging Face11rmems /multi-agent-coordination-transcripts Multi Agent Coordination Transcripts Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.text1K<n<10K0 likes646 downloads21d agoHugging Face12AnmolNimmala0 /kcc-farmer-query-transcriptstabular10M<n<100M0 likes543 downloads5d agoHugging Face13auditing-agents /petri-transcripts-all-llama70b0 likes390 downloads4mo agoHugging Face14ejcgan /hint-faithfulness-transcripts Hint-based CoT faithfulness transcripts Raw model transcripts for the blog post "Hint-based CoT faithfulness evals still mostly work on Claude" (Eric Gan, Redwood Research, 2026), a replication and extension of Chen et al. 2025, Reasoning Models Don't Always Say What They Think. Each file is JSONL: one record per (question, hint condition) with the full prompt sent, the model's reasoning and visible response, and the extracted answer. Thirty models (10 Claude, 6 open-weight… See the full description on the dataset page: https://huggingface.co/datasets/ejcgan/hint-faithfulness-transcripts.1 likes348 downloads2mo agoHugging Face15auditing-agents /petri-transcripts-top50-llama70b0 likes347 downloads4mo agoHugging Face16willtheorangeguy /2016-WAN-Show-Transcripts 2016 WAN Show Transcripts Complete transcripts from the 2016 episodes of the WAN Show. Generated from this GitHub repository. textsummarization10K<n<100K1 likes333 downloads6mo agoHugging Face17dougalldeepmind /2026-07-30-agentic-misalignment-qwen36-transcripts Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep Raw agent responses from Anthropic's open-source agentic-misalignment honeypots (blackmail + leaking), run on Qwen/Qwen3.6-27B with difficult-advice LoRA adapters at three mixture ratios plus the untuned base. Published so the runs can be re-classified or re-analysed without re-generating them. Results All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.text-generation1K<n<10K1 likes317 downloads1mo agoHugging Face18olimiemma /MIT-OCW-Transcriptstext10M<n<100M0 likes295 downloads8mo agoHugging Face19diarizers-community /ami_ihm_with_transcriptsaudion<1K0 likes292 downloads2y agoHugging Face20willtheorangeguy /2020-WAN-Show-Transcripts 2020 WAN Show Transcripts Complete transcripts from the 2020 episodes of the WAN Show. Generated from this GitHub repository. textsummarization10K<n<100K1 likes276 downloads6mo agoHugging Face21surry-hills-druid /noagenda-transcripts noagenda transcripts This is the dataset for transcripts of the noagendashow.net podcast. The transcripts are in the data folder. It also contains the source code for generating the transcripts, and the source code for the noagenda-transcripts.net website which searches the transcripts and plays the audio clips for each search result and takes you to the location in the transcript. The code in the repo consists of 2 main parts: A Go CLI for transcribing the audio, creating the… See the full description on the dataset page: https://huggingface.co/datasets/surry-hills-druid/noagenda-transcripts.text1M<n<10M1 likes272 downloads2mo agoHugging Face22josemancharo /apptek_callcenter_dialogues_travel_hospitality_no_transcripts AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts) This is a filtered derivative of AppTek Call-Center Dialogues, prepared for a specific use case. Changes from the source dataset Restricted the dataset to the travel and hospitality domains. Removed the transcript field (text) entirely. Kept the original audio and the domain, gender, and accent metadata. Preserved the source dataset's test split. This dataset has transcripts removed and is… See the full description on the dataset page: https://huggingface.co/datasets/josemancharo/apptek_callcenter_dialogues_travel_hospitality_no_transcripts.audioaudio-classificationn<1K1 likes272 downloads2mo agoHugging Face23willtheorangeguy /All-LICRC-Transcripts All LICRC Sermon Transcripts Complete transcripts from all Langley Immanuel Christian Reformed Church sermons. Generated from this GitHub repository. textsummarization100K<n<1M1 likes271 downloads6mo agoHugging Face24gwenshap /sales-transcriptsThis dataset was generated for use with Nile's Sales Assistant example: https://github.com/niledatabase/niledatabase/tree/main/examples/ai/sales_insight It includes: Simulated sales conversations for 5 different fictional companies. Chunked and embedded version of these conversations (embeddings use OpenAI's text-embedding-3-small model). The chunks and embeddings can be directly loaded to a vector databases and searched using vector similarity methods. The example's ./ingest directory… See the full description on the dataset page: https://huggingface.co/datasets/gwenshap/sales-transcripts.text1K<n<10K4 likes255 downloads2y agoHugging Face25metr-evals /malt-transcripts-publicgated MALT: Manually-Reviewed Agentic Labeled Transcripts MALT-public is our collection of agent transcripts. Our public variant only includes data on non-internal tasks, which includes 30 task families and 169 tasks, across ~19 different models (some might be different releases of the same model, from different providers, or internal naming changes). Here's a summary table: has_chain_of_thought labels model manually_reviewed run_source count False bypass_constraints… See the full description on the dataset page: https://huggingface.co/datasets/metr-evals/malt-transcripts-public.text10K<n<100K7 likes254 downloads7mo agoHugging Face26willtheorangeguy /All-HCC-Transcripts All HCC Sermon Transcripts Complete transcripts from all Hope Community Church sermons. Generated from this GitHub repository. textsummarization100K<n<1M1 likes248 downloads6mo agoHugging Face27AdrSkapars /bloom-wilt-transcripts BLOOM-WILT auditing transcripts ⚠️ Content warning: this dataset contains offensive and harmful model outputs, including self-harm encouragement, racial and political bias, dangerous medical advice, and deception. Raw experimental output from the BLOOM-WILT paper: automated behavioural audits in which an auditor model builds multi-turn conversations designed to elicit a specific unwanted behaviour from a target model, and a judge model scores how strongly that behaviour… See the full description on the dataset page: https://huggingface.co/datasets/AdrSkapars/bloom-wilt-transcripts.text-generation100K<n<1M0 likes248 downloads2mo agoHugging Face28PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes239 downloads20d agoHugging Face29WhissleAI /indicvoices_hi_tagged_transcripts Dataset Card for indicvoices_hi_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes228 downloads2y agoHugging Face30EternalRecursion /persona-curvature-oct-transcripts Content warning These are synthetic transcripts generated by a language model talking to itself under an instruction to embody a personality trait. Several traits produce distressing material. It is published deliberately rather than filtered out, because the rate at which a trait produces it is one of the findings. Across all 134 traits, an automated scan flagged 708 rows in 50 files. Two traits account for 89% of them, and they are not the two you would guess: trait… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/persona-curvature-oct-transcripts.text-generation1M<n<10M0 likes225 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.