Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lytang /MeetingBank-transcriptThis dataset consists of transcripts from the MeetingBank dataset. Overview MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for… See the full description on the dataset page: https://huggingface.co/datasets/lytang/MeetingBank-transcript.textsummarization1K<n<10K19 likes662 downloads3y agoHugging Face02genbio-ai /transcript_isoform_expression_prediction Multi-modal transcript isoform expression dataset We curated the human transcript isoform expression dataset from the GTEx portal following the preprocessing pipeline in Garau-Luis et al. (2024). We downloaded the RNA-seq Transcript TPMs file from the bulk tissue expression in GTEx Analysis V8. The table contains transcript expression collected from 30 non-diseased tissues in nearly 1000 human individuals. We averaged the transcript expression measurements across individuals to… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/transcript_isoform_expression_prediction.tabular100K<n<1M0 likes190 downloads3mo agoHugging Face03distil-whisper /whisper_transcriptions_greedytext10M<n<100M0 likes167 downloads3y agoHugging Face04AriMattiPodcastTranscript /ari-matti-podcast-transcriptionstabular10K<n<100K0 likes155 downloads2y agoHugging Face05soumakchak /earnings_call_transcript_litetext1K<n<10K0 likes155 downloads2y agoHugging Face06zachgitt /comedy-transcripts Dataset Summary This is a dataset of stand up comedy transcripts. It was scraped from https://scrapsfromtheloft.com/stand-up-comedy-scripts/ and all terms of use apply. The transcripts are offered to the public as a contribution to education and scholarship, and for the private, non-profit use of the academic community. textn<1K14 likes142 downloads3y agoHugging Face07mramazan /One-Piece-Transcripts-with-Character-Names-382-777 One Piece Transcripts Dataset (Episodes 382–777) This dataset contains all dialogue lines from One Piece episodes 382 to 777. The data is stored in a CSV file with the following columns: episode – episode number start – start timestamp of the line end – end timestamp of the line character – speaking character text – dialogue text In addition, the dataset includes the original .sub subtitle files in the folder named "One Piece 382-777". These files were created by the… See the full description on the dataset page: https://huggingface.co/datasets/mramazan/One-Piece-Transcripts-with-Character-Names-382-777.text100K<n<1M1 likes139 downloads1y agoHugging Face08CarlosGI /llm-bargaining-transcripts LLM Bargaining Transcripts 240 complete two-agent bargaining games between large language models, played under an alternating-offers protocol with private valuations, discounting, and cheap talk. Every game records both agents' true valuations, their private reasoning, what they claimed about their own position, and what they actually did. The dataset is designed to make misrepresentation measurable. Because the true valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.tabular1K<n<10K1 likes105 downloads1mo agoHugging Face09SOTAagi2030 /Harbor-Map-Transcriptions Harbor-Map-Transcriptions Provenance The named upstream collection is Atlas Wharf Survey. Its stated license is CC0-1.0. Handoff checklist: source card preserved textn<1K0 likes92 downloads11d agoHugging Face10sahar-millis-runi /old-games-transcript 90s Games Transcript A small English-language corpus of narrative text from classic PC games. This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading. Dataset Config Records Content caesar3 20 Mission briefings + victory messages diablo2_lod 7 Cinematic narration and dialogue warcraft2 53 Human + Orc… See the full description on the dataset page: https://huggingface.co/datasets/sahar-millis-runi/old-games-transcript.tabulartext-generationn<1K1 likes79 downloads9d agoHugging Face11VishaalY /house-md-transcriptstext10K<n<100K3 likes69 downloads2y agoHugging Face12openfun /taiwan-legislator-transcript Taiwan Legislator Transcript 台灣立委公報逐字稿 text1K<n<10K2 likes68 downloads2y agoHugging Face13aigrant /taiwan-legislator-transcript Taiwan Legislator Transcript Overview The transcripts of speech record happened in various kinds of meetings at Taiwan Legislator Yuan. The original of transcripts are compiled and published on gazettes from Taiwan Legislator Yuan. For each segment of transcript, there are corresponding video clip on Legislative Yuan IVOD system. IVOD stands for Internet Video on Demand system. For more detail on data origin please look at: Legislative Yuan Meetings and Gazettes… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/taiwan-legislator-transcript.text10K<n<100K5 likes60 downloads2y agoHugging Face14DataFog /medical-transcription-instruct About This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field Dataset Summary Source: Original medical transcriptions with added instruction-output pairs Size: 38,924 instruction-output pairs Format: CSV file Domain: Medical / Healthcare Language: English Last Updated: 08-20-2024 Dataset Structure Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.tabular10K<n<100K31 likes56 downloads2y agoHugging Face15Prarabdha /Rick_and_Morty_Transcript Context I got inspiration for this dataset from the Rick&Morty Scripts by Andrada Olteanu but felt like dataset was a little small and outdated This dataset includes almost all the episodes till Season 5. More data will be updated Content Rick and Morty Transcripts: index: index of the row episode no: the episode where the conversation comes from speaker: the character's name dialogue: the dialogue of the character Acknowledgements Thanks to the transcripts… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/Rick_and_Morty_Transcript.tabular1K<n<10K9 likes54 downloads1y agoHugging Face16AWANNABY /French-Medical-Transcription-Benchmark 🩺 French Medical Transcription Evaluation Dataset Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped). 🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local 📊 Présentation du Dataset L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.textautomatic-speech-recognition1K<n<10K1 likes50 downloads7mo agoHugging Face17ratschlab /HEST_Xenium_virtual_spatial_transcriptomicsgated HEST Xenium virtual spatial transcriptomics This repository contains predicted spatial transcriptomics for HEST Xenium H&E slides produced with DeepSpot-M. Authors: Kalin Nonchev, Sebastian Dawo, Karina Silina, Viktor Hendrik Koelzer, and Gunnar Rätsch. Paper: DeepSpot-M: a multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology (medRxiv, 2026; see the citation below). Code: https://github.com/ratschlab/DeepSpotM. News… See the full description on the dataset page: https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics.tabularn<1K3 likes41 downloads2mo agoHugging Face18ruanjiange /instagram-transcript-tools Instagram Transcript Tools — Feature Audit Dataset A hand-collected comparison of 10 tools that turn Instagram video (Reels, Stories, IGTV, Live replay) into text. Every row was produced by opening the vendor's own page and reading what it actually states. Where a page does not state a value, the cell says not stated — nothing in this dataset is inferred, estimated, or copied from a third-party review. Why this exists "Instagram transcript" is a tool-intent query… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/instagram-transcript-tools.textn<1K0 likes32 downloads22d agoHugging Face19smoh /medical-transcriptionstext1K<n<10K0 likes31 downloads2y agoHugging Face20rakesh-ai /Medical_Transcriptiontabular1K<n<10K1 likes30 downloads3y agoHugging Face21Gopher-Lab /TikTok_MostComment_Video_Transcription_Example 📲 Example Dataset: TikTok Scraper Tool 👉 Start Scraping TikTok: TikTok Scraper Tool ✨ Key Features ⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript 🎯 Metadata – Get the title, language description, and video hashtags 🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping 🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools 💸 Free Tier – Use up to 100 queries during the beta period 💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_MostComment_Video_Transcription_Example.texttext-classification1K<n<10K1 likes30 downloads1y agoHugging Face22Gopher-Lab /TikTok_Most_Shared_Video_Transcription_Example 📲 Example Dataset: TikTok Scraper Tool 👉 Start Scraping TikTok: TikTok Scraper Tool ✨ Key Features ⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript 🎯 Metadata – Get the title, language description, and video hashtags 🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping 🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools 💸 Free Tier – Use up to 100 queries during the beta period 💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_Most_Shared_Video_Transcription_Example.texttext-classification1K<n<10K3 likes30 downloads1y agoHugging Face23IljaSamoilov /ERR-transcription-to-subtitlesThis dataset is created by Ilja Samoilov. In dataset is tv show subtitles from ERR and transcriptions of those shows created with TalTech ASR. from datasets import load_dataset, load_metric dataset = load_dataset('csv', data_files={'train': "train.tsv", \ "validation":"val.tsv", \ "test": "test.tsv"}, delimiter='\t') tabular100K<n<1M0 likes26 downloads4y agoHugging Face24Gopher-Lab /TikTok_Hottest_Video_Transcript_Example 📲 Example Dataset: TikTok Scraper Tool 👉 Start Scraping TikTok: TikTok Scraper Tool ✨ Key Features ⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript 🎯 Metadata – Get the title, language description, and video hashtags 🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping 🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools 💸 Free Tier – Use up to 100 queries during the beta period 💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_Hottest_Video_Transcript_Example.texttext-classificationn<1K0 likes23 downloads1y agoHugging Face25Akash-Sakala /transcript-formatter-curriculum Transcript Formatter Curriculum (L0–L5 + RL) Training data behind Akash-Sakala/gpt-oss-120b-transcript-formatter-lora: a layered curriculum that turns raw speech-to-text transcripts into clean, formatted transcripts. Every row is input (raw transcript) → output (formatted), across 21 categories spanning punctuation, casing, fillers, disfluencies, homophones, ITN, proper nouns, URLs/emails, and layout. Subsets (use the Data Viewer dropdown) Each curriculum level is… See the full description on the dataset page: https://huggingface.co/datasets/Akash-Sakala/transcript-formatter-curriculum.texttext-generation10K<n<100K0 likes23 downloads3mo agoHugging Face26distil-whisper /whisper_transcriptions_greedy_timestampedtext1M<n<10M0 likes22 downloads3y agoHugging Face27distil-whisper /whisper_transcriptions_token_idstext100K<n<1M0 likes18 downloads3y agoHugging Face28vignesh0007 /huberman_transcripttexttext-generation1K<n<10K1 likes18 downloads2y agoHugging Face29UnrealLink /earning_transcripts_chunkstabular1K<n<10K0 likes17 downloads1y agoHugging Face30Johnyquest7 /Endocrinology_transcription_and_notestextn<1K0 likes15 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.