datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parliament_hearings_processed
Preprocessed parliament hearings ASR dataset to truecased form.
Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126
dataset_info:
features:
- name: id
dtype: string
- name: audio
dtype:
audio:
sampling_rate: 16000
- name: transcription
sequence: string
splits:
- name: train
num_bytes: 53645064353.18
num_examples: 191455
- name: test
num_bytes: 740331298.0
num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.malaysia-parliament-billstw_parliament_split
Parliament
Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the
zh_tw split of disco-eth/WorldSpeech.
It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD
parliamentary proceedings.
This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style
text filtering pipeline. Audio is preserved from the upstream dataset and cast
as a Hugging Face Audio(sampling_rate=24000) feature.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.swg_parliament_fhnwparliament
Parliament
Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the
zh_tw split of disco-eth/WorldSpeech.
It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD
parliamentary proceedings.
This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style
text filtering pipeline. Audio is preserved from the upstream dataset and cast
as a Hugging Face Audio(sampling_rate=24000) feature.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/parliament.canadian-parliamentary-expenditures
Canadian House of Commons Parliamentary Expenditures Dataset
This dataset contains detailed expenditure records from the Canadian House of Commons, spanning from 2021 Q2 to 2025 Q4, with 1,219,648 total expenditure records across 450 parliament members.
Dataset Structure
parliamentary_data_hf/
├── data/
│ ├── train/ # Training split (2021-2024)
│ │ ├── expenditures-2021-q2.parquet
│ │ ├── expenditures-2021-q3.parquet
│ │ ├── ...… See the full description on the dataset page: https://huggingface.co/datasets/irf23/canadian-parliamentary-expenditures.swiss_parliament_corpus
Dataset Card for "swiss_parliament_corpus"
More Information needed
turkish_parliamentary_data
Grand National Assembly Corpus of Türkiye (GNACT)
A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts.
Loading the dataset
from datasets import load_dataset
# Strategy 1: full session documents, all bodies (default)
ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train")
#… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.norwegian_parliament
Dataset Card Creation Guide
Dataset Summary
This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.malaysia-parliament-youtube
Malaysia Parliament Youtube
Entire videos from https://www.youtube.com/@PARLIMENMALAYSIA1 including Live.
With total 2068 audio files, total 3985.4 hours.
how to download
huggingface-cli download --repo-type dataset \
--include '*.z*' \
--local-dir './' \
malaysia-ai/malaysia-parliament-youtube
wget https://www.7-zip.org/a/7z2301-linux-x64.tar.xz
tar -xf 7z2301-linux-x64.tar.xz
~/7zz x parlimen.zip -y -mmt40
Source code
Source code at… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysia-parliament-youtube.basque_parliament_1
Dataset Card for Basque Parliament Speech Corpus 1.0
This work was partially funded by the Spanish Ministry of Science and Innovation (OPENSPEECH
project, PID2019-106424RB-I00).
Dataset Summary
The Basque Parliament Speech Corpus 1.0 consists of 1462 hours of speech extracted from
Basque Parliament plenary sessions from 2013 to 2022. Encoded as MP3 files, the dataset
contains 759192 transcribed segments either spoken in Basque, Spanish or both (in
Basque and Spanish).… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/basque_parliament_1.speeches-of-the-german-parliamenthansard-embeddings
Hansard Embeddings & Derived Analytics
Text embeddings and derived analytics cubes covering the parliamentary
Hansard of four Australian jurisdictions: the Commonwealth (au), New
South Wales (nsw), South Australia (sa), and Western Australia (wa).
No Hansard prose is in this dataset. It ships vectors, join keys, and
derived metadata (subject headings, member names, parties, counts). The
official record remains with each parliament; to hydrate the text, run the
open-source… See the full description on the dataset page: https://huggingface.co/datasets/parliament-data/hansard-embeddings.norwegian_parliamentde_corpora_parliament_processedparliamentrag-camera-leg19
ParliamentRAG — Italian Chamber of Deputies, 19th legislature
Full proceedings of the Italian Chamber of Deputies (Camera dei deputati) for the 19th legislature, from the first sitting on 13 October 2022 through 6 August 2026: verbatim speech transcripts, roll-call votes with every individual ballot, parliamentary acts with EuroVoc subjects, and the deputies' group, committee and government memberships over time. The dataset is refreshed as new sittings are ingested.
The tables… See the full description on the dataset page: https://huggingface.co/datasets/emeierkeio/parliamentrag-camera-leg19.dutch_corpora_parliament_processeden_corpora_parliament_processedfr_corpora_parliament_processedsv_corpora_parliament_processednl_corpora_parliament_processedsv_corpora_parliament_processedes_corpora_parliament_processedsv_corpora_parliament_processed_v0es_corpora_parliament_processedgerman-parliament-speeches
German Parliament Speeches
This dataset contains speeches from the German parliament, derived from the Open Discourse Project (Harvard Dataverse).
Source
Data source:
Open Discourse ProjectHarvard DataverseDOI: 10.7910/DVN/FIKIBO
Original citation:
@data{DVN/FIKIBO_2020,
author = {Richter, Florian and Koch, Philipp and Franke, Oliver and Kraus, Jakob and Kuruc, Fabrizio and Thiem, Anja and Högerl, Judith and Heine, Stella and Schöps, Konstantin},
publisher = {Harvard… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/german-parliament-speeches.bulgarian-parliament-moss-1to1
Bulgarian Parliament MOSS 1-to-1
Private machine-labelled training candidates under construction. Not gold data
or a held-out evaluation set. Public redistribution terms remain unverified.
Source: DimitarV/eurospeech-bg-single-speaker at
6aa43432891867bcd249c0fe944ac6286888231b, derived from
disco-eth/EuroSpeech.
Reference transcripts are parliamentary stenographic text. Speaker identities
are inferred clusters, not named or human-verified speakers.
Active batch list… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-parliament-moss-1to1.uk-parliament-hansard-modern
uk-parliament-hansard-modern (FineWeb-style)
Modern UK Hansard debates exported into LLM-training-friendly FineWeb-style Parquet shards.
What’s inside
Format: nanochat-parquet-v1
Layout: shard_*.parquet + metadata.json
Text column: text
Parquet settings: zstd (level 3), row_group_size=1024, use_dictionary=False, write_statistics=False
HuggingFace
Dataset repo: JayJayThrowThrow/uk-parliament-hansard-modern
Loading
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/JayJayThrowThrow/uk-parliament-hansard-modern.Hellenic-greek-parliamentary-speech
HParl: Hellenic Parliamentary Speech Corpus
Dataset Description
Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors.
Link to the original source: https://inventory.clarin.gr/corpus/1602
HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.HansardNER
HansardNER
HansardNER is a named entity recognition (NER) dataset for Sinhala parliamentary proceedings. It contains 4,000 speaker turns (1.49 million words) from 494 sitting days of the Sri Lankan Hansard, 2017–2026, labelled with nine entity classes.
The dataset has two kinds of labels:
Silver (all 4,000 turns): labelled by a large language model (Google Gemini) under written guidelines. Not checked by a person.
Gold (648 of those turns): the silver labels corrected by human… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/HansardNER.
