Team Ai
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NickWeng /tw_parliament_split Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.audioautomatic-speech-recognition100K<n<1M0 likes1.1k downloads23d agoHugging Face02OpenFormosa /parliament Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/parliament.audioautomatic-speech-recognition100K<n<1M2 likes674 downloads4mo agoHugging Face03alimetin /turkish-parliament-speechaudioautomatic-speech-recognition1K<n<10K1 likes70 downloads9mo agoHugging Face04stcoats /au-parliament-poc-turns-2026-03-30 Australian Parliamentary Speech Corpus — proof of concept This dataset is a proof-of-concept time-aligned speech corpus derived from the Australian House of Representatives sitting of 30 March 2026. It contains 531 speaker-attributed turns with embedded 16 kHz audio, automatic speech recognition transcripts, source timestamps, parliamentary metadata, and selected speaker and constituency metadata. The source recording is ParlView recording 4454892. The resource demonstrates a… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/au-parliament-poc-turns-2026-03-30.audioautomatic-speech-recognitionn<1K0 likes33 downloads12d agoHugging Face05intronhealth /afrispeech-parliamentgated AfriSpeech-Parliament: Transcribed Parliamentary Sessions from Four African Nations This work is licensed under aCreative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. Overview AfriSpeech-Parliament is a 35.86-hour dataset of transcribed parliamentary sessions from Nigeria, Ghana, South Africa, and Kenya. Sourced from their respective official YouTube channels, each sample is a 16-second audio chunk with corresponding manual transcription.… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrispeech-parliament.audioautomatic-speech-recognition1K<n<10K1 likes13 downloads1y agoHugging Face06deepdml /basque_parliament_1gatedaudioautomatic-speech-recognition1M<n<10M0 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.