Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TCLResearchEurope /intelligent_wakeup Intelligent Wakeup A synthetic corpus for device-directed speech detection: multi-speaker conversations in which most speech is not addressed to the voice assistant, with the moments that are clearly marked. Conventional assistants detect a wake word but cannot tell whether what follows is meant for them. This corpus is built to train and evaluate the module that makes that decision from the whole session, not from an isolated command. Project page:… See the full description on the dataset page: https://huggingface.co/datasets/TCLResearchEurope/intelligent_wakeup.audioaudio-classification1K<n<10K0 likes407 downloads28d agoHugging Face02kizuna-intelligence /AItuber-Persona-Voices-JA AItuber Persona Voices JA 195体のAITuberペルソナに対し、キャラクター設定に基づいた声質設計・セリフ生成・音声合成を行ったデータセットです。 概要 項目 値 ペルソナ数 195 総音声ファイル数 20,800 (参照音声195 + 発話20,600) 音声フォーマット WAV, PCM 16-bit, 44.1kHz, mono セリフカテゴリ original, descriptive, emotional 言語 日本語 データ構造 各行は1つの音声ファイルに対応し、以下のカラムを持ちます: カラム 型 説明 persona_id string ペルソナ識別子 (persona_000 ~ persona_194) persona_index int ペルソナインデックス(元データセットの行位置に対応) persona_name string キャラクター名 voice_description_ja… See the full description on the dataset page: https://huggingface.co/datasets/kizuna-intelligence/AItuber-Persona-Voices-JA.audiotext-to-speech10K<n<100K6 likes220 downloads6mo agoHugging Face03InteliLab /ict_s2s_refactoredaudio100K<n<1M0 likes174 downloads4mo agoHugging Face04IntelLabs /Intel_Robotic_Welding_Multimodal_Datasetgated Dataset Card for the Intel Robotic Welding Multimodal Dataset This dataset was collected to enable multimodal welding defect detection research. The dataset contains over 4000 annotated samples and was collected in an automotive production floor setting in collaboration with a supplier with access to such facilities. Each sample contains a video, associated audio, a time-series from welding sensors, and five post-weld images for a particular weld. A separately licensed… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/Intel_Robotic_Welding_Multimodal_Dataset.audio67 likes173 downloads1y agoHugging Face05kizuna-intelligence /gemini-tan-ita Geminiたん Geminiたんのイメージボイスです。Kizuna Intelligence株式会社の声質生成技術で作りました。 同じ声で読み上げた日本語の合成音声424本(約40分)です。 音声・台詞・ファイル名をParquetにまとめています。下の一覧で試聴でき、Filesからダウンロードできます。 台詞とファイル名の対応は transcripts.csv でも確認できます。ダウンロード数はこのページに表示されます。 全件の聴取確認は未実施です。 ライセンスは Apache 2.0。元のITAコーパスの台詞は Unlicense です。 audiotext-to-speechn<1K0 likes118 downloads19d agoHugging Face06intelsense /indicvoices-bengali Dataset Description This dataset has been collected from IndicVoices-R github repository. Only the Bengali language portion is in this dataset. Train Duration: 393985.0155 seconds Test Duration: 2719.5667 seconds audiotext-to-speech10K<n<100K0 likes48 downloads2y agoHugging Face07intelsense /openslr-banglaaudio1K<n<10K2 likes40 downloads2y agoHugging Face08NIO-LLM /IntelliASR-Benchgated Training audio dataset 41,390 examples and 41,390 WAV audio files. Extract audio.zip beside train_with_metadata_and_tts.jsonl. The archive contains only audio files with no subdirectories. JSONL wav values are relative filenames. Other fields and audio bytes are unchanged. Only filesystem paths were anonymized; audio and transcripts were not redacted. unzip audio.zip SHA-256: e42261a45cccdfa2398791993f04511841970ef8705aec08d88c5f0dc663d054 audio.zip… See the full description on the dataset page: https://huggingface.co/datasets/NIO-LLM/IntelliASR-Bench.audio10K<n<100K0 likes33 downloads15d agoHugging Face09intelmedica /medical-tts-parquet-2-16khzgated IntelMedica Medical TTS Dataset v2 (16kHz) Description Synthetic medical speech dataset for fine-tuning Whisper-based ASR models on clinical and nursing terminology. Contains 101,475 audio-text pairs totaling 184.1 hours of speech at 16 kHz mono, generated using Kokoro-82M TTS with 19 voices across three English accent groups. This is v2 -- a companion to the v1 dataset (125,500 samples, ~257 hours). v2 focuses on terms from additional data sources (RxNorm API, FDA… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-2-16khz.audioautomatic-speech-recognition100K<n<1M3 likes28 downloads6mo agoHugging Face10intelmedica /medical-tts-parquet-1gated IntelMedica Medical TTS Dataset v1 (24kHz) -- DEPRECATED This dataset is deprecated. Please use intelmedica/medical-tts-parquet-1-16khz instead, which contains 125,500 samples (vs 10,000 here) at 16kHz sample rate optimized for ASR training. Description Synthetic medical speech dataset for training medical ASR models. This is the original 10K-sample version at 24kHz. It has been superseded by the 16kHz version with 12.5x more data. Dataset Details Samples:… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-1.audioautomatic-speech-recognition10K<n<100K0 likes16 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.