Team Ai
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MrSupW /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K38 likes1.4k downloads1y agoHugging Face02lilonghao /MM-ContextASR-Bench MM-ContextASR Bench Metadata and evaluation splits for Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark. Dataset summary Config Examples Audio Context Primary metric mm_contextasr 1,250 (250 current utterances × 5 histories) 1,439 WAV files included Controlled user-assistant dialogue entity Recall kespeech 19,212 Source ID only Same-speaker speech and transcript CER, SER, entity Recall cv_yue 3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.audioautomatic-speech-recognition10K<n<100K1 likes306 downloads24d agoHugging Face03wayu-ai /thai-contextasr-bench Thai Contextual-Biasing ASR Benchmark TL;DR Does your Thai ASR system actually use the context you give it (e.g., a list of names, custom words from your own dictionary) — and does it hallucinate when the context is irrelevant? Each utterance comes with a bias list: entity strings (brands, person names, places) that may or may not be spoken in the audio, written the way a real Thai user would write them — one list, mixed Thai and Latin script. A good system does… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-contextasr-bench.audioautomatic-speech-recognition1K<n<10K1 likes228 downloads2mo agoHugging Face04maikezu /asr-context-induced-leakage When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR Overview SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/asr-context-induced-leakage.audioautomatic-speech-recognition1K<n<10K0 likes67 downloads5mo agoHugging Face05bsmu666 /ContextASR-Bench ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/bsmu666/ContextASR-Bench.textautomatic-speech-recognition10K<n<100K0 likes58 downloads9mo agoHugging Face06Ethan615 /taiwan-conversation-context-100-domainsgated Taiwan Conversation Context 100 Domains Dataset Description Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。 本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。 資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於: 語音生成資料前處理 Text-to-Speech, TTS Spoken Dialogue Generation Conversational AI Customer Service Dialogue Modeling Role-play Dialogue Dataset 台灣繁體中文語音模型訓練 生活情境問答模型訓練 對話式 AI 助理訓練 RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.texttext-generation1M<n<10M2 likes42 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.