datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tw_parliament_split
Parliament
Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the
zh_tw split of disco-eth/WorldSpeech.
It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD
parliamentary proceedings.
This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style
text filtering pipeline. Audio is preserved from the upstream dataset and cast
as a Hugging Face Audio(sampling_rate=24000) feature.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.parliament
Parliament
Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the
zh_tw split of disco-eth/WorldSpeech.
It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD
parliamentary proceedings.
This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style
text filtering pipeline. Audio is preserved from the upstream dataset and cast
as a Hugging Face Audio(sampling_rate=24000) feature.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/parliament.turkish-parliament-speechau-parliament-poc-turns-2026-03-30
Australian Parliamentary Speech Corpus — proof of concept
This dataset is a proof-of-concept time-aligned speech corpus derived from the
Australian House of Representatives sitting of 30 March 2026. It contains
531 speaker-attributed turns with embedded 16 kHz audio, automatic speech
recognition transcripts, source timestamps, parliamentary metadata, and selected
speaker and constituency metadata. The source recording is ParlView recording
4454892.
The resource demonstrates a… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/au-parliament-poc-turns-2026-03-30.afrispeech-parliament
AfriSpeech-Parliament: Transcribed Parliamentary Sessions from Four African Nations
This work is licensed under aCreative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Overview
AfriSpeech-Parliament is a 35.86-hour dataset of transcribed parliamentary sessions from Nigeria, Ghana, South Africa, and Kenya. Sourced from their respective official YouTube channels, each sample is a 16-second audio chunk with corresponding manual transcription.… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrispeech-parliament.basque_parliament_1
