Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.1k downloads4y agoHugging Face02Teejeigh /raw_friends_series_transcriptRaw transcript from friends tv series, chunk into ~1000 token lines. Text tagged by character. language: - en size_categories: - n<1K textn<1K2 likes720 downloads3y agoHugging Face03rmems /multi-agent-coordination-transcripts Multi Agent Coordination Transcripts Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.text1K<n<10K0 likes662 downloads23d agoHugging Face04Deltarunefan /Deltarune-Complete-Transcript-Cleaned Deltarune Chapters 1–4 Dataset Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora. Why This Exists As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite their… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.texttext-generation10K<n<100K4 likes210 downloads7mo agoHugging Face05yeong-hwan /2024-earnings-call-transcripttabular1K<n<10K0 likes187 downloads2y agoHugging Face06shantanugoel /aawaaz-transcript-cleanup-dataset Aawaaz Transcript Cleanup Dataset Training pairs for cleaning messy speech transcripts (ASR output, voice dictation) into well-formatted text while preserving the speaker's voice and meaning. Dataset Description Each example is a pair of: input: A realistic messy transcript with filler words, false starts, self-corrections, grammar errors, and missing punctuation output: The cleaned version with fillers removed, grammar fixed, punctuation added, and domain-appropriate… See the full description on the dataset page: https://huggingface.co/datasets/shantanugoel/aawaaz-transcript-cleanup-dataset.texttext-generation10K<n<100K0 likes148 downloads7mo agoHugging Face07lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes122 downloads1mo agoHugging Face08lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes113 downloads1mo agoHugging Face09SOTAagi2030 /Beacon-Transcript-Curated Harbor Light Museum — Curated Transcripts This card contains reviewed lighthouse transcript records prepared by the Harbor Light Museum. Curatorial status: provenance packet published Provenance ledger record_id collection cataloged_on verification page_count BL-203 Lantern Room Logs 2024-11-18 verified 21 BL-117 Fog Signal Registers 2024-10-03 verified 36 BL-088 Keeper Correspondence 2024-06-22 verified 18 BL-241 Harbor Weather Sheets 2024-03-15… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Beacon-Transcript-Curated.textn<1K0 likes110 downloads14d agoHugging Face10iblai /ibl-khanacademy-transcripts ibl-khanacademy-transcripts This dataset houses the transcripts of openly available videos from Khan Academy. The transcripts were scrapped from Khan Academy's youtube channel text1K<n<10K2 likes86 downloads3y agoHugging Face11tomngdev /shell-safety-transcriptsConverted from tomngdev/shell-safety into conversations transcripts. Structure is for my own training with static system prompt and changing <SessionContext> block texttext-classification10K<n<100K0 likes80 downloads2mo agoHugging Face12hsanyyasyn97gmail /transcripts-separatetextn<1K0 likes76 downloads1mo agoHugging Face13AltaySec /altayduel-transcripts 🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2) Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir. 📌 TL;DR 2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu. 439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.tabulartext-classification1K<n<10K1 likes75 downloads2mo agoHugging Face14podongchip /119-call-transcript-synthetic-labelsgated 119 Call Transcript (Synthetic) + Patient Labels 한국 응급실 ↔ 119 구급대원 / 병원 간 통화를 모사한 합성(synthetic) 한국어 통화 전사문 13,382건과, 통화마다 환자 수·환자별 성별·나이대·Pre-KTAS 주증상 라벨입니다. "통화에서 환자 정보를 추출해 주변 병원에 일괄 전달하는 모델"의 학습·평가용으로 만들었습니다. A synthetic Korean dataset of 13,382 emergency-call transcripts (EMS ↔ ED, or hospital ↔ hospital), each labeled with the number of patients and, per patient, sex, age band, and a Pre-KTAS (Korean prehospital triage) chief complaint. All records are generated text; no… See the full description on the dataset page: https://huggingface.co/datasets/podongchip/119-call-transcript-synthetic-labels.texttext-classification10K<n<100K0 likes71 downloads10h agoHugging Face15spikecodes /911-call-transcriptstextn<1K3 likes69 downloads2y agoHugging Face16Noothi /huberman-lab-transcripts Huberman Lab Transcript Dataset Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel. Dataset 438 videos 9,833 transcript chunks ~114 million characters JSONL format Each record contains: ext ideo_id itle Processing The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.texttext-generation1K<n<10K0 likes59 downloads1mo agoHugging Face17shuyuej /CC-BY-STEMM-Podcast-Transcriptstext10K<n<100K1 likes53 downloads2y agoHugging Face18SOTAagi2030 /field-notes-transcription-packet Field-notes Transcription Packet Verified field-note transcriptions retained for the research archive. Retained notes: 6 Featured note: NOTE-1002 — Cedar / English Transcription window: 2023-01-18T16:20:00Z to 2023-12-15T10:00:00Z Site counts (Cedar/Lark/Morrow/Pine): 2/2/1/1 Mean word counts (Cedar/Lark/Morrow/Pine): 422.5/362.5/360.0/480.0 textn<1K0 likes52 downloads1mo agoHugging Face19dsvv-cair /ipa-transcription-datase 🗣️ English Text → IPA Transcription Dataset Overview This dataset provides a large-scale, phonemically rich collection of English text paired with International Phonetic Alphabet (IPA) transcriptions, designed to support research and applications in speech-language pathology, phonetics, and natural language processing. It was created to enable data-driven phonetic transcription, reducing reliance on traditional rule-based systems and supporting modern… See the full description on the dataset page: https://huggingface.co/datasets/dsvv-cair/ipa-transcription-datase.texttext-generation100K<n<1M0 likes50 downloads5mo agoHugging Face20courtnoski /Farsight-SRV-Transcripts The Farsight Institute: Scientific Remote Viewing (SRV) Transcripts Dataset Summary This dataset contains the complete, unabridged archive of Scientific Remote Viewing (SRV) session transcripts and project summaries produced by The Farsight Institute, directed by Dr. Courtney Brown. The data consists of hundreds of highly detailed, text-rich transcripts describing historical events, planetary mysteries, and extraterrestrial dynamics. All remote viewing sessions… See the full description on the dataset page: https://huggingface.co/datasets/courtnoski/Farsight-SRV-Transcripts.texttext-generationn<1K0 likes49 downloads4mo agoHugging Face21Infektyd /council-transcripts Council Multi-Agent Deliberation Transcripts Real deliberation transcripts from Council, a multi-agent orchestration skill for OpenClaw. What This Is Council routes tasks to specialized Grok persona agents — Workhorse (deep technical reasoning), Creative (novel ideas, chaos energy), and Speed (fast iteration) — then synthesizes their outputs into a final verdict via a Conductor. These transcripts capture the full deliberation process: prompts, per-persona responses, and… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/council-transcripts.tabulartext-generationn<1K0 likes41 downloads6mo agoHugging Face22ontocord /MixtureVitae-finevideo-transcripts-onlyThis is the whisper transcription from the Finvideo dataset. There are artifacts esp when music or sound is mistakenly transcribed as words. TBD: cleanup these repetitious words. text10K<n<100K0 likes37 downloads1y agoHugging Face23ReopenAI /cantonese-youtube-transcription-fusionhttps://huggingface.co/datasets/alvanlii/cantonese-youtube 数据集中train-00000-of-01090.parquet 到 train-00350-of-01090.parquet 部分的转写文本清洗。使用qwen3-asr、qwen3-omni、sensevoicesmall(https://huggingface.co/ASLP-lab/WSYue-ASR) 进行转写,然后用 Qwen3.6-35B-A3B 根据语义进行转写纠正。 text100K<n<1M0 likes33 downloads2mo agoHugging Face24Dbmaxwell /pufi-duf-transcriptstabularn<1K0 likes33 downloads28d agoHugging Face25bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes32 downloads5mo agoHugging Face26hsanyyasyn97gmail /transcripts-mergedtextn<1K0 likes32 downloads1mo agoHugging Face27Whispering-GPT /whisper-transcripts-the-vergeannotations_creators: machine-generated language: en language_creators: crowdsourced license: [] multilinguality: monolingual paperswithcode_id: wikitext-2 pretty_name: Whisper-Transcripts size_categories: 1M<n<10M source_datasets: original tags: [] task_categories: text-generation fill-mask task_ids: language-modeling masked-language-modeling text1K<n<10K4 likes29 downloads4y agoHugging Face28bingbangboom /cleaned-asr-transcriptstexttext-generation10K<n<100K1 likes29 downloads7mo agoHugging Face29aarongrainer /atc-audio-transcriptstext1K<n<10K1 likes28 downloads1y agoHugging Face30AWANNABY /luciolescribe-transcription-faq 🎙️ LucioleScribe Transcription FAQ - Dataset Production v2.0 📋 Description Dataset enrichi de questions-réponses FAQ sur la transcription IA 100% locale avec LucioleScribe, première plateforme française de transcription conforme RGPD par conception. 🎯 Caractéristiques clés 📊 Taille: 182+ paires question-réponse (expansion continue) 🏷️ Métadonnées: Enrichi avec catégories, intents, buyer stages, styles 🌍 Langue: Français (France) 📑 Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/luciolescribe-transcription-faq.textquestion-answeringn<1K0 likes27 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.