datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.synthetic-speaker-diarization-dataset-fa-3000tamil-english-podcast-diarization
Tamil-English Code-Mixed Podcast Diarization Dataset
Dataset Summary
This dataset contains long-form Tamil-English code-mixed podcast recordings
annotated for speaker diarization research. The recordings consist of natural
conversational speech with multiple speakers and realistic acoustic conditions,
making the dataset suitable for evaluating diarization pipelines in
real-world scenarios.
The dataset is intended to support research in:
Speaker diarization
Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/Rangasuthan/tamil-english-podcast-diarization.synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.danish-diarization-bench
Danish Diarization Benchmark (Synthetic) — v2
A 3996-row synthetic speaker-diarization benchmark in Danish, built by mixing
single-speaker utterances from
syvai/danish-asr-unified
into multi-speaker recordings.
What changed in v2 (2026-05-18)
Per-segment text — each entry in segments now carries its text field directly. The redundant parallel texts column has been removed. Old consumers that joined segments[i] with texts[i] should switch to segments[i]["text"].
Silent… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-diarization-bench.sada-diarization-preview
SADA 2022 Arabic Diarization
Training-ready speaker-attributed ASR windows derived from
SADA 2022. The source
recordings are mirrored at
khaledalganem/sada2022.
Splits
train: 36,004 windows, 202.064 hours, 4,062 recordings
validation: 853 windows, 4.774 hours, 88 recordings
test: 901 windows, 5.006 hours, 111 recordings
Total: 37,758 windows and
211.844 hours.
The official SADA train, validation, and test partitions are preserved.
Windows are 8–28 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Mohaddz/sada-diarization-preview.mac-m4pro-fresh-diarization-demucs-20260902
gdrive-sbpn-fresh-diarization-demucs-mac-m4pro-20260902
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/mac-m4pro-fresh-diarization-demucs-20260902.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 1,536
Chunk audio duration: 5.773 hours
Source transcript rows represented: 4,326
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
22
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-05.gdrive-sbpn-supplied-diarization-demucs-terminal-l4-20260814
gdrive-sbpn-supplied-diarization-demucs-terminal-l4-20260814
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-supplied-diarization-demucs-terminal-l4-20260814.gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814
gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01
This dataset contains lossless FLAC chunks derived from 45 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 1,034
Chunk audio duration: 4.411 hours
Source transcript rows represented: 3,041
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
9
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-01.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 1,097
Chunk audio duration: 5.143 hours
Source transcript rows represented: 3,875
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
19
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-02.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 2,116
Chunk audio duration: 6.624 hours
Source transcript rows represented: 5,601
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
56
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-06.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 650
Chunk audio duration: 5.893 hours
Source transcript rows represented: 4,213
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
1
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-03.gdrive-sbpn-fresh-diarization-demucs-h100-20260815
gdrive-sbpn-fresh-diarization-demucs-h100-20260815
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-demucs-h100-20260815.gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04
gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04
This dataset contains lossless FLAC chunks derived from 50 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 599
Chunk audio duration: 4.567 hours
Source transcript rows represented: 2,961
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
0
Music annotation… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-colab-l4-20260813-archive-04.gdrive-sbpn-diarization-before-after-review-20260812
gdrive-sbpn-diarization-before-after-review-20260812
Each row is one transcript row whose timestamp changed during diarization boundary correction. Play audio_before_diarization and audio_after_diarization side by side. Both are cut from the same original recording; Demucs is not used. Text and speaker are unchanged.
start_movement_seconds and end_movement_seconds equal after minus before: negative means earlier and positive means later. boundary_audit_json records the fusion… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-diarization-before-after-review-20260812.
