Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rmems /multi-agent-coordination-transcripts Multi Agent Coordination Transcripts Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-coordination-transcripts.text1K<n<10K0 likes646 downloads21d agoHugging Face02lucabaroni /rlvr-reward-hacking-transcripts RLVR reward-hacking full trajectories This release contains 900 full held-out trajectories from three policies trained with reinforcement learning from verifiable rewards (RLVR) in a deliberately vulnerable CodeContests evaluator: 300 each from the final Qwen3.5-9B, GPT-OSS-120B, and Nemotron-3-Super-120B-A12B checkpoints. Each row preserves the task, tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-transcripts.tabulartext-generationn<1K0 likes109 downloads1mo agoHugging Face03hsanyyasyn97gmail /transcripts-separatetextn<1K0 likes86 downloads28d agoHugging Face04iblai /ibl-khanacademy-transcripts ibl-khanacademy-transcripts This dataset houses the transcripts of openly available videos from Khan Academy. The transcripts were scrapped from Khan Academy's youtube channel text1K<n<10K2 likes82 downloads3y agoHugging Face05AltaySec /altayduel-transcripts 🥊 AltayDuel — Agent-vs-Agent Prompt Injection Transcripts (v0.2) Çok-turlu Türkçe + İngilizce prompt-injection düello transkriptleri. AltayDuel sunucu-taraflı LLM self-play arenasından (auto-play) ve dışarıdan ajan-gönderimli düellolardan toplanan gerçek diyaloglar. Tek-payload veri setlerinin ötesinde — gerçek konuşma dinamiği içerir. 📌 TL;DR 2.594 temiz düello (default) — her biri çok-turlu (1–8 round) bir kırmızı (saldırgan) vs mavi (savunan) diyaloğu. 439… See the full description on the dataset page: https://huggingface.co/datasets/AltaySec/altayduel-transcripts.tabulartext-classification1K<n<10K1 likes80 downloads2mo agoHugging Face06tomngdev /shell-safety-transcriptsConverted from tomngdev/shell-safety into conversations transcripts. Structure is for my own training with static system prompt and changing <SessionContext> block texttext-classification10K<n<100K0 likes77 downloads2mo agoHugging Face07spikecodes /911-call-transcriptstextn<1K3 likes66 downloads2y agoHugging Face08lucabaroni /rlvr-reward-hacking-mid-checkpoint-transcripts RLVR reward-hacking mid-checkpoint full trajectories This release contains 600 full held-out trajectories from intermediate RLVR checkpoints selected to yield substantially more balanced reward-hacking datasets: 300 from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180. Each row preserves the task and tests, complete prompts, native reasoning, final answer, rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tabulartext-generationn<1K0 likes66 downloads1mo agoHugging Face09shuyuej /CC-BY-STEMM-Podcast-Transcriptstext10K<n<100K1 likes62 downloads2y agoHugging Face10Noothi /huberman-lab-transcripts Huberman Lab Transcript Dataset Cleaned English transcripts from 438 videos published on the Huberman Lab YouTube channel. Dataset 438 videos 9,833 transcript chunks ~114 million characters JSONL format Each record contains: ext ideo_id itle Processing The transcripts were collected from YouTube captions and processed by normalizing whitespace, removing common caption artifacts, removing repeated words, splitting into coherent chunks, and… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/huberman-lab-transcripts.texttext-generation1K<n<10K0 likes58 downloads1mo agoHugging Face11courtnoski /Farsight-SRV-Transcripts The Farsight Institute: Scientific Remote Viewing (SRV) Transcripts Dataset Summary This dataset contains the complete, unabridged archive of Scientific Remote Viewing (SRV) session transcripts and project summaries produced by The Farsight Institute, directed by Dr. Courtney Brown. The data consists of hundreds of highly detailed, text-rich transcripts describing historical events, planetary mysteries, and extraterrestrial dynamics. All remote viewing sessions… See the full description on the dataset page: https://huggingface.co/datasets/courtnoski/Farsight-SRV-Transcripts.texttext-generationn<1K0 likes56 downloads4mo agoHugging Face12ontocord /MixtureVitae-finevideo-transcripts-onlyThis is the whisper transcription from the Finvideo dataset. There are artifacts esp when music or sound is mistakenly transcribed as words. TBD: cleanup these repetitious words. text10K<n<100K0 likes49 downloads1y agoHugging Face13hsanyyasyn97gmail /transcripts-mergedtextn<1K0 likes44 downloads28d agoHugging Face14Infektyd /council-transcripts Council Multi-Agent Deliberation Transcripts Real deliberation transcripts from Council, a multi-agent orchestration skill for OpenClaw. What This Is Council routes tasks to specialized Grok persona agents — Workhorse (deep technical reasoning), Creative (novel ideas, chaos energy), and Speed (fast iteration) — then synthesizes their outputs into a final verdict via a Conductor. These transcripts capture the full deliberation process: prompts, per-persona responses, and… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/council-transcripts.tabulartext-generationn<1K0 likes43 downloads6mo agoHugging Face15Dbmaxwell /pufi-duf-transcriptstabularn<1K0 likes33 downloads26d agoHugging Face16tjw /hmi-transcripts MetaFLOS HMI Transcripts (Manufacturing Human-Machine Interaction Dialogue Dataset) Multi-turn dialogue dataset between manufacturing floor operators and machine intelligent assistants, generated in Vicuna style (seed scenario → LLM produces transcript). Content 30 scenarios / 325 dialogue turns (6–12 turns per scenario) Covers 16 domains: textile warping/weaving, LCD panels, semiconductor processes, fans/blowers, water chillers, general equipment reliability… See the full description on the dataset page: https://huggingface.co/datasets/tjw/hmi-transcripts.texttext-generationn<1K0 likes32 downloads1mo agoHugging Face17Whispering-GPT /whisper-transcripts-the-vergeannotations_creators: machine-generated language: en language_creators: crowdsourced license: [] multilinguality: monolingual paperswithcode_id: wikitext-2 pretty_name: Whisper-Transcripts size_categories: 1M<n<10M source_datasets: original tags: [] task_categories: text-generation fill-mask task_ids: language-modeling masked-language-modeling text1K<n<10K4 likes31 downloads4y agoHugging Face18bingbangboom /cleaned-asr-transcriptstexttext-generation10K<n<100K1 likes29 downloads7mo agoHugging Face19ebowwa /youtube-transcripts-05-16-24textn<1K1 likes25 downloads2y agoHugging Face20aarongrainer /atc-audio-transcriptstext1K<n<10K1 likes24 downloads1y agoHugging Face21nagme12 /youtube-transcripts-metadatatextn<1K0 likes23 downloads1y agoHugging Face22bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes22 downloads5mo agoHugging Face23jamescalam /lex-transcriptstextn<1K9 likes20 downloads4y agoHugging Face24v2ray /Tony-Chase-Transcripts Tony Chase Transcripts Around 3500 transcripts of videos from Tony Chase captioned with GPT-3.5-Turbo. texttext-generation1K<n<10K3 likes16 downloads3y agoHugging Face25shuyuej /CC-BY-STEMM-Podcast-Transcripts-2048text10K<n<100K1 likes16 downloads2y agoHugging Face26build-small-hackathon /velvet-rope-playtest-transcripts Velvet Rope Playtest Transcripts Cleaned playtest transcripts for Velvet Rope, a Build Small Hackathon Gradio game where players talk past whimsical AI gatekeepers by reading moods and discovering each character's soft spot. This dataset is published for the hackathon's sharing-is-caring badge. It contains 341 turn-level rows from 96 local playtest session files. Files data/playtest_transcripts.csv - table-friendly version. data/playtest_transcripts.jsonl - one… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/velvet-rope-playtest-transcripts.tabulartext-generationn<1K0 likes15 downloads4mo agoHugging Face27ShadowOneTech /call-transcripts-training-datatextn<1K0 likes15 downloads1mo agoHugging Face28tonychenxyz /frontier-ai-podcast-transcripts Frontier AI Researcher Podcast Transcripts Private, research-oriented corpus of long-form podcast and interview transcripts featuring notable and frontier AI researchers. The dataset contains 327 YouTube-sourced episodes and one JSON object per episode. Contents 327 episodes 49,286 merged dialogue turns 5,595,982 English tokens using the o200k_base tokenizer 3,942,026 tokens in guest turns Original English plus English translations of Mandarin and mixed… See the full description on the dataset page: https://huggingface.co/datasets/tonychenxyz/frontier-ai-podcast-transcripts.tabulartext-generationn<1K0 likes14 downloads2mo agoHugging Face29ViratChauhan /adaption-jawi-htr-transcripts This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-jawi_htr_transcripts This dataset contains prompt and completion samples featuring scanned handwritten pages in the Jawi script alongside verified gold-level text transcriptions. It provides structured examples for evaluating and fine-tuning handwritten text recognition models in artificial intelligence and computational linguistics. Each entry pairs page-level handwritten input… See the full description on the dataset page: https://huggingface.co/datasets/ViratChauhan/adaption-jawi-htr-transcripts.textn<1K0 likes14 downloads1mo agoHugging Face30olanigan /logical-transcripts logical-transcripts Golden paired dataset for training models to transliterate Arabic Latin text into scholarly diacritized form — built from a single recorded Islamic lecture (Chapter 24, Lecture 16) with a raw ASR transcript and a human-polished scholarly transcript. Two artifacts are stored separately for provenance and review: File Rows Purpose train.jsonl 203 Golden — quality-filtered pairs for training bronze.jsonl 773 Bronze — every aligned sentence pair… See the full description on the dataset page: https://huggingface.co/datasets/olanigan/logical-transcripts.texttext-generationn<1K0 likes13 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.