Team Ai
14 results

parliamentary

irf23 /canadian-parliamentary-expenditures Canadian House of Commons Parliamentary Expenditures Dataset This dataset contains detailed expenditure records from the Canadian House of Commons, spanning from 2021 Q2 to 2025 Q4, with 1,219,648 total expenditure records across 450 parliament members. Dataset Structure parliamentary_data_hf/ ├── data/ │ ├── train/ # Training split (2021-2024) │ │ ├── expenditures-2021-q2.parquet │ │ ├── expenditures-2021-q3.parquet │ │ ├── ...… See the full description on the dataset page: https://huggingface.co/datasets/irf23/canadian-parliamentary-expenditures.tabulartabular-classification1M<n<10M0 likes640 downloads1y agoHugging Faceboun-tabilab /turkish_parliamentary_data Grand National Assembly Corpus of Türkiye (GNACT) A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts. Loading the dataset from datasets import load_dataset # Strategy 1: full session documents, all bodies (default) ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train") #… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.tabulartext-generation1M<n<10M8 likes473 downloads6mo agoHugging FaceElormiden /Hellenic-greek-parliamentary-speech HParl: Hellenic Parliamentary Speech Corpus Dataset Description Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors. Link to the original source: https://inventory.clarin.gr/corpus/1602 HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.audio10K<n<100K1 likes98 downloads1y agoHugging Facesl-parliamentary-nlp /HansardNER HansardNER HansardNER is a named entity recognition (NER) dataset for Sinhala parliamentary proceedings. It contains 4,000 speaker turns (1.49 million words) from 494 sitting days of the Sri Lankan Hansard, 2017–2026, labelled with nine entity classes. The dataset has two kinds of labels: Silver (all 4,000 turns): labelled by a large language model (Google Gemini) under written guidelines. Not checked by a person. Gold (648 of those turns): the silver labels corrected by human… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/HansardNER.texttoken-classification1K<n<10K0 likes68 downloads8d agoHugging Facesl-parliamentary-nlp /sl-parliamentary-hansard-17-26 Dataset Card for Sri Lanka Parliamentary Hansard Sri Lanka Parliamentary Hansard is a trilingual parliamentary speech corpus built from publicly available Hansard records of the Parliament of Sri Lanka. It contains Sinhala (සිංහල), Tamil (தமிழ்), English, and code-mixed speeches from 2017 to 2026, with speaker names, dates, and topic-modeling labels. The dataset was created for the research paper "Trilingual Topic Modeling of Sri Lankan Parliamentary Debates", associated with… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/sl-parliamentary-hansard-17-26.tabular10K<n<100K2 likes57 downloads1mo agoHugging FaceDecima-Data /luxembourgish-parliamentary-corpus Luxembourgish Parliamentary Corpus (2023–2028) A provenance-documented, speaker-attributed corpus of Luxembourg's parliamentary proceedings, built from the official session reports (comptes rendus / "D'Chamberblietchen") of the Chambre des Députés, legislature 2023–2028. Luxembourgish (Lëtzebuergesch) is a documented low-resource language: the Luxembourgish Wikipedia holds roughly 64,000 articles and most large language models perform poorly in it for lack of training material.… See the full description on the dataset page: https://huggingface.co/datasets/Decima-Data/luxembourgish-parliamentary-corpus.texttext-generation100K<n<1M0 likes28 downloads3mo agoHugging Face