datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
norwegian_parliament
Dataset Card Creation Guide
Dataset Summary
This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.HansardNER
HansardNER
HansardNER is a named entity recognition (NER) dataset for Sinhala parliamentary proceedings. It contains 4,000 speaker turns (1.49 million words) from 494 sitting days of the Sri Lankan Hansard, 2017–2026, labelled with nine entity classes.
The dataset has two kinds of labels:
Silver (all 4,000 turns): labelled by a large language model (Google Gemini) under written guidelines. Not checked by a person.
Gold (648 of those turns): the silver labels corrected by human… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/HansardNER.mk-parliament-qa
Macedonian Parliamentary QA (RAG Evaluation)
A question-answering benchmark in Macedonian, built over transcripts of plenary
sessions of the Assembly of the Republic of North Macedonia. It is intended for
evaluating retrieval-augmented generation systems on a low-resource language,
with questions that require single-passage, multi-passage, and cross-document
reasoning.
Dataset structure
Two configs are provided.
questions
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/stojchevamarija/mk-parliament-qa.Polish-Parliamentary-Speeches-Corpus
Dataset Card for Polish Parliamentary Speeches Corpus (PPSC)
Dataset Description
The Polish Parliamentary Speeches Corpus (PPSC) is a collection of official transcripts of parliamentary speeches made by Polish politicians. It was created to facilitate the modeling of political viewpoints in a low-resource language (Polish) using supervised fine-tuning. The dataset assigns speeches to binary ideological categories (Left-wing and Right-wing) based on the speakers'… See the full description on the dataset page: https://huggingface.co/datasets/PoliWings/Polish-Parliamentary-Speeches-Corpus.luxembourgish-parliamentary-corpus
Luxembourgish Parliamentary Corpus (2023–2028)
A provenance-documented, speaker-attributed corpus of Luxembourg's parliamentary
proceedings, built from the official session reports (comptes rendus /
"D'Chamberblietchen") of the Chambre des Députés, legislature 2023–2028.
Luxembourgish (Lëtzebuergesch) is a documented low-resource language: the
Luxembourgish Wikipedia holds roughly 64,000 articles and most large language
models perform poorly in it for lack of training material.… See the full description on the dataset page: https://huggingface.co/datasets/Decima-Data/luxembourgish-parliamentary-corpus.parliamentary-debate-casesnation-parliamentary-prompts
