datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
intelligent_wakeup
Intelligent Wakeup
A synthetic corpus for device-directed speech detection: multi-speaker conversations in
which most speech is not addressed to the voice assistant, with the moments that are
clearly marked.
Conventional assistants detect a wake word but cannot tell whether what follows is meant for
them. This corpus is built to train and evaluate the module that makes that decision from the
whole session, not from an isolated command.
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/TCLResearchEurope/intelligent_wakeup.AItuber-Persona-Voices-JA
AItuber Persona Voices JA
195体のAITuberペルソナに対し、キャラクター設定に基づいた声質設計・セリフ生成・音声合成を行ったデータセットです。
概要
項目
値
ペルソナ数
195
総音声ファイル数
20,800 (参照音声195 + 発話20,600)
音声フォーマット
WAV, PCM 16-bit, 44.1kHz, mono
セリフカテゴリ
original, descriptive, emotional
言語
日本語
データ構造
各行は1つの音声ファイルに対応し、以下のカラムを持ちます:
カラム
型
説明
persona_id
string
ペルソナ識別子 (persona_000 ~ persona_194)
persona_index
int
ペルソナインデックス(元データセットの行位置に対応)
persona_name
string
キャラクター名
voice_description_ja… See the full description on the dataset page: https://huggingface.co/datasets/kizuna-intelligence/AItuber-Persona-Voices-JA.ict_s2s_refactoredIntel_Robotic_Welding_Multimodal_Dataset
Dataset Card for the Intel Robotic Welding Multimodal Dataset
This dataset was collected to enable multimodal welding defect detection research. The dataset contains over 4000 annotated samples and was collected in an automotive production floor setting in collaboration with a supplier with access to such facilities. Each sample contains a video, associated audio, a time-series from welding sensors, and five post-weld images for a particular weld. A separately licensed… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/Intel_Robotic_Welding_Multimodal_Dataset.gemini-tan-ita
Geminiたん
Geminiたんのイメージボイスです。Kizuna Intelligence株式会社の声質生成技術で作りました。
同じ声で読み上げた日本語の合成音声424本(約40分)です。
音声・台詞・ファイル名をParquetにまとめています。下の一覧で試聴でき、Filesからダウンロードできます。
台詞とファイル名の対応は transcripts.csv でも確認できます。ダウンロード数はこのページに表示されます。
全件の聴取確認は未実施です。
ライセンスは Apache 2.0。元のITAコーパスの台詞は Unlicense です。
indicvoices-bengali
Dataset Description
This dataset has been collected from IndicVoices-R github repository. Only the Bengali language portion is in this dataset.
Train Duration: 393985.0155 seconds
Test Duration: 2719.5667 seconds
openslr-banglaIntelliASR-Bench
Training audio dataset
41,390 examples and 41,390 WAV audio files.
Extract audio.zip beside train_with_metadata_and_tts.jsonl. The archive contains only audio files with no subdirectories. JSONL wav values are relative filenames. Other fields and audio bytes are unchanged.
Only filesystem paths were anonymized; audio and transcripts were not redacted.
unzip audio.zip
SHA-256:
e42261a45cccdfa2398791993f04511841970ef8705aec08d88c5f0dc663d054 audio.zip… See the full description on the dataset page: https://huggingface.co/datasets/NIO-LLM/IntelliASR-Bench.medical-tts-parquet-2-16khz
IntelMedica Medical TTS Dataset v2 (16kHz)
Description
Synthetic medical speech dataset for fine-tuning Whisper-based ASR models on clinical and nursing terminology. Contains 101,475 audio-text pairs totaling 184.1 hours of speech at 16 kHz mono, generated using Kokoro-82M TTS with 19 voices across three English accent groups.
This is v2 -- a companion to the v1 dataset (125,500 samples, ~257 hours). v2 focuses on terms from additional data sources (RxNorm API, FDA… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-2-16khz.medical-tts-parquet-1
IntelMedica Medical TTS Dataset v1 (24kHz) -- DEPRECATED
This dataset is deprecated. Please use intelmedica/medical-tts-parquet-1-16khz instead, which contains 125,500 samples (vs 10,000 here) at 16kHz sample rate optimized for ASR training.
Description
Synthetic medical speech dataset for training medical ASR models. This is the original 10K-sample version at 24kHz. It has been superseded by the 16kHz version with 12.5x more data.
Dataset Details
Samples:… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-1.
