diarization
bengali-diarization-synthetic-v3voicevox-diarization-ja
VOICEVOX合成 多話者日本語音声(話者分離検証用データセット)
Qwen3-ASR の話者分離ファインチューニング検証、特に embedding層+projectorのみの学習で話者分離が学べるか を確かめるために作成した、正解ラベルが厳密な合成データです。
作り方
gemma(gemma-4-31B-it)で話題別のモノローグ日本語テキストを生成
文単位に分割し、各文境界で確率0.3で話者を切替(2〜3話者)
各文を割当てた VOICEVOX 話者で音声合成し連結
→ 合成由来なので、各文の [開始秒, 終了秒] と話者が完全に既知(人手や別ASRに依存しない厳密ラベル)。
規模
698クリップ、各 約2分、opus (24kbps) / 16kHz 相当 mono
2話者 約88% / 3話者 約12%
話者は各クリップ内の登場順に spk_0, spk_1, spk_2 と採番
列 (data.jsonl)
列
内容… See the full description on the dataset page: https://huggingface.co/datasets/okadahiroaki/voicevox-diarization-ja.synthetic-speaker-diarization-dataset-fa-large-3000Speaker-Diarization-Instructions
Speaker-Diarization-Instructions
Convert diarization dataset from https://huggingface.co/diarizers-community into speech instructions dataset and chunk max to 30 seconds because most of speech encoder use for LLM come from Whisper Encoder.
We highly recommend to not include AMI test set from both AMI-IHM and AMI-SDM in training set to prevent contamination. This dataset supposely to become a speech diarization benchmark.
how to prepare the dataset
huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Speaker-Diarization-Instructions.shemo_diarization_datasetluganda_callhome_diarization_dataset_MHDP
