Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jkot /parliament_hearings_processed Preprocessed parliament hearings ASR dataset to truecased form. Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126 dataset_info: features: - name: id dtype: string - name: audio dtype: audio: sampling_rate: 16000 - name: transcription sequence: string splits: - name: train num_bytes: 53645064353.18 num_examples: 191455 - name: test num_bytes: 740331298.0 num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.audio100K<n<1M1 likes16k downloads3y agoHugging Face02NickWeng /tw_parliament_split Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.audioautomatic-speech-recognition100K<n<1M0 likes1.1k downloads23d agoHugging Face03OpenFormosa /parliament Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/parliament.audioautomatic-speech-recognition100K<n<1M2 likes674 downloads4mo agoHugging Face04yanickschraner /swiss_parliament_corpus Dataset Card for "swiss_parliament_corpus" More Information needed audio10K<n<100K1 likes430 downloads4y agoHugging Face05Elormiden /Hellenic-greek-parliamentary-speech HParl: Hellenic Parliamentary Speech Corpus Dataset Description Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors. Link to the original source: https://inventory.clarin.gr/corpus/1602 HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.audio10K<n<100K2 likes134 downloads1y agoHugging Face06alimetin /turkish-parliament-speechaudioautomatic-speech-recognition1K<n<10K1 likes70 downloads9mo agoHugging Face07malaysia-ai /pseudolabel-malaysia-parliament-youtube-whisper-large-v3 pseudolabel-malaysia-parliament-youtube-whisper-large-v3 Pseudolabel malaysia-ai/malaysia-parliament-youtube using openai/whisper-large-v3 How to prepare the dataset huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \ malaysia-ai/pseudolabel-malaysia-parliament-youtube-whisper-large-v3 wget… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-malaysia-parliament-youtube-whisper-large-v3.audio100K<n<1M0 likes39 downloads1y agoHugging Face08stcoats /au-parliament-poc-turns-2026-03-30 Australian Parliamentary Speech Corpus — proof of concept This dataset is a proof-of-concept time-aligned speech corpus derived from the Australian House of Representatives sitting of 30 March 2026. It contains 531 speaker-attributed turns with embedded 16 kHz audio, automatic speech recognition transcripts, source timestamps, parliamentary metadata, and selected speaker and constituency metadata. The source recording is ParlView recording 4454892. The resource demonstrates a… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/au-parliament-poc-turns-2026-03-30.audioautomatic-speech-recognitionn<1K0 likes33 downloads12d agoHugging Face09huseinzol05 /VC-Parliamentaudio0 likes32 downloads1y agoHugging Face10juanhebert /id_corpora_parliament_processedaudio1K<n<10K0 likes14 downloads4y agoHugging Face11intronhealth /afrispeech-parliamentgated AfriSpeech-Parliament: Transcribed Parliamentary Sessions from Four African Nations This work is licensed under aCreative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. Overview AfriSpeech-Parliament is a 35.86-hour dataset of transcribed parliamentary sessions from Nigeria, Ghana, South Africa, and Kenya. Sourced from their respective official YouTube channels, each sample is a 16-second audio chunk with corresponding manual transcription.… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrispeech-parliament.audioautomatic-speech-recognition1K<n<10K1 likes13 downloads1y agoHugging Face12cpersia /parliament_14263audio10K<n<100K0 likes12 downloads8mo agoHugging Face13shiimi /dv_corpora_parliament_processedaudio1K<n<10K0 likes10 downloads1y agoHugging Face14deepdml /basque_parliament_1gatedaudioautomatic-speech-recognition1M<n<10M0 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.