datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.raw_friends_series_transcriptRaw transcript from friends tv series, chunk into ~1000 token lines. Text tagged by character.
language:
- en
size_categories:
- n<1K
multi-agent-coordination-transcripts
Multi Agent Coordination Transcripts
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.Deltarune-Complete-Transcript-Cleaned
Deltarune Chapters 1–4 Dataset
Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora.
Why This Exists
As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite their… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.2024-earnings-call-transcriptaawaaz-transcript-cleanup-dataset
Aawaaz Transcript Cleanup Dataset
Training pairs for cleaning messy speech transcripts (ASR output, voice dictation) into well-formatted text while preserving the speaker's voice and meaning.
Dataset Description
Each example is a pair of:
input: A realistic messy transcript with filler words, false starts, self-corrections, grammar errors, and missing punctuation
output: The cleaned version with fillers removed, grammar fixed, punctuation added, and domain-appropriate… See the full description on the dataset page: https://huggingface.co/datasets/shantanugoel/aawaaz-transcript-cleanup-dataset.rlvr-reward-hacking-transcripts
RLVR reward-hacking full trajectories
This release contains 900 full held-out trajectories from three policies trained with
reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable
CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B
checkpoints. Each row preserves the task, tests, complete prompts, native
reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.Beacon-Transcript-Curated
Harbor Light Museum — Curated Transcripts
This card contains reviewed lighthouse transcript records prepared by the Harbor Light Museum.
Curatorial status: provenance packet published
Provenance ledger
record_id
collection
cataloged_on
verification
page_count
BL-203
Lantern Room Logs
2024-11-18
verified
21
BL-117
Fog Signal Registers
2024-10-03
verified
36
BL-088
Keeper Correspondence
2024-06-22
verified
18
BL-241
Harbor Weather Sheets
2024-03-15… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Beacon-Transcript-Curated.transcripts-separateibl-khanacademy-transcripts
ibl-khanacademy-transcripts
This dataset houses the transcripts of openly available videos from Khan Academy.
The transcripts were scrapped from Khan Academy's youtube channel
altayduel-transcripts
🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2)
Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir.
📌 TL;DR
2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu.
439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.shell-safety-transcriptsConverted from tomngdev/shell-safety into conversations transcripts.
Structure is for my own training with static system prompt and changing <SessionContext> block
911-call-transcriptsrlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.CC-BY-STEMM-Podcast-Transcriptslogistics-cx-transcript-analysis-chatml
OmniCX Logistics CX Dataset (Research Preview)
Dataset Summary
This dataset is designed for structured extraction of logistics and customer-experience (CX) signals from multi-turn support conversations.
Each record uses ChatML-style messages with:
a fixed system instruction
a user transcript
an assistant JSON payload matching LogisticsCXMetrics
This release is a research preview and should not be treated as a production-certified benchmark.
Project repository:… See the full description on the dataset page: https://huggingface.co/datasets/mangesh-ux/logistics-cx-transcript-analysis-chatml.huberman-lab-transcripts
Huberman Lab Transcript Dataset
Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel.
Dataset
438 videos
9,833 transcript chunks
~114 million characters
JSONL format
Each record contains:
ext
ideo_id
itle
Processing
The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.Farsight-SRV-Transcripts
The Farsight Institute: Scientific Remote Viewing (SRV) Transcripts
Dataset Summary
This dataset contains the complete, unabridged archive of Scientific Remote Viewing (SRV) session transcripts and project summaries produced by The Farsight Institute, directed by Dr. Courtney Brown.
The data consists of hundreds of highly detailed, text-rich transcripts describing historical events, planetary mysteries, and extraterrestrial dynamics. All remote viewing sessions… See the full description on the dataset page: https://huggingface.co/datasets/courtnoski/Farsight-SRV-Transcripts.field-notes-transcription-packet
Field-notes Transcription Packet
Verified field-note transcriptions retained for the research archive.
Retained notes: 6
Featured note: NOTE-1002 — Cedar / English
Transcription window: 2023-01-18T16:20:00Z to 2023-12-15T10:00:00Z
Site counts (Cedar/Lark/Morrow/Pine): 2/2/1/1
Mean word counts (Cedar/Lark/Morrow/Pine): 422.5/362.5/360.0/480.0
119-call-transcript-synthetic-labels
119 Call Transcript Synthetic Labels
한국 119 응급 신고 통화를 바탕으로 만든 합성(synthetic) 텍스트에, Claude Haiku 4.5로 구조화 라벨링(연령/성별/의식/활력징후/증상/KTAS 등급 등)을 붙인 데이터셋입니다.
This dataset contains synthetic Korean 119 (emergency call) transcripts, labeled with structured fields (age, sex, consciousness, vital signs, symptoms, KTAS triage level, etc.) using Claude Haiku 4.5.
데이터 구성 / Structure
train.jsonl: 7,196 rows
val.jsonl: 1,794 rows
전체 8,990건은 모두 합성(synthetic) 통화이며, 실제 통화 STT 기반… See the full description on the dataset page: https://huggingface.co/datasets/podongchip/119-call-transcript-synthetic-labels.MixtureVitae-finevideo-transcripts-onlyThis is the whisper transcription from the Finvideo dataset. There are artifacts esp when music or sound is mistakenly transcribed as words. TBD: cleanup these repetitious words.
ipa-transcription-datase
🗣️ English Text → IPA Transcription Dataset
Overview
This dataset provides a large-scale, phonemically rich collection of English text paired with International Phonetic Alphabet (IPA) transcriptions, designed to support research and applications in speech-language pathology, phonetics, and natural language processing.
It was created to enable data-driven phonetic transcription, reducing reliance on traditional rule-based systems and supporting modern… See the full description on the dataset page: https://huggingface.co/datasets/dsvv-cair/ipa-transcription-datase.transcripts-mergedcouncil-transcripts
Council Multi-Agent Deliberation Transcripts
Real deliberation transcripts from Council, a multi-agent orchestration skill for OpenClaw.
What This Is
Council routes tasks to specialized Grok persona agents — Workhorse (deep technical reasoning), Creative (novel ideas, chaos energy), and Speed (fast iteration) — then synthesizes their outputs into a final verdict via a Conductor. These transcripts capture the full deliberation process: prompts, per-persona responses, and… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/council-transcripts.yt-transcriptionscantonese-youtube-transcription-fusionhttps://huggingface.co/datasets/alvanlii/cantonese-youtube 数据集中train-00000-of-01090.parquet 到 train-00350-of-01090.parquet 部分的转写文本清洗。使用qwen3-asr、qwen3-omni、sensevoicesmall(https://huggingface.co/ASLP-lab/WSYue-ASR)
进行转写,然后用 Qwen3.6-35B-A3B 根据语义进行转写纠正。
pufi-duf-transcriptshmi-transcripts
MetaFLOS HMI Transcripts (Manufacturing Human-Machine Interaction Dialogue Dataset)
Multi-turn dialogue dataset between manufacturing floor operators and machine intelligent assistants, generated in Vicuna style (seed scenario → LLM produces transcript).
Content
30 scenarios / 325 dialogue turns (6–12 turns per scenario)
Covers 16 domains: textile warping/weaving, LCD panels, semiconductor processes, fans/blowers, water chillers, general equipment reliability… See the full description on the dataset page: https://huggingface.co/datasets/tjw/hmi-transcripts.whisper-transcripts-the-vergeannotations_creators:
machine-generated
language:
en
language_creators:
crowdsourced
license: []
multilinguality:
monolingual
paperswithcode_id: wikitext-2
pretty_name: Whisper-Transcripts
size_categories:
1M<n<10M
source_datasets:
original
tags: []
task_categories:
text-generation
fill-mask
task_ids:
language-modeling
masked-language-modeling
luciolescribe-transcription-faq
🎙️ LucioleScribe Transcription FAQ - Dataset Production v2.0
📋 Description
Dataset enrichi de questions-réponses FAQ sur la transcription IA 100% locale avec LucioleScribe, première plateforme française de transcription conforme RGPD par conception.
🎯 Caractéristiques clés
📊 Taille: 182+ paires question-réponse (expansion continue)
🏷️ Métadonnées: Enrichi avec catégories, intents, buyer stages, styles
🌍 Langue: Français (France)
📑 Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/luciolescribe-transcription-faq.
