Team Ai
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sl-parliamentary-nlp /HansardNER HansardNER HansardNER is a named entity recognition (NER) dataset for Sinhala parliamentary proceedings. It contains 4,000 speaker turns (1.49 million words) from 494 sitting days of the Sri Lankan Hansard, 2017–2026, labelled with nine entity classes. The dataset has two kinds of labels: Silver (all 4,000 turns): labelled by a large language model (Google Gemini) under written guidelines. Not checked by a person. Gold (648 of those turns): the silver labels corrected by human… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/HansardNER.texttoken-classification1K<n<10K0 likes73 downloads12d agoHugging Face02PoliWings /Polish-Parliamentary-Speeches-Corpus Dataset Card for Polish Parliamentary Speeches Corpus (PPSC) Dataset Description The Polish Parliamentary Speeches Corpus (PPSC) is a collection of official transcripts of parliamentary speeches made by Polish politicians. It was created to facilitate the modeling of political viewpoints in a low-resource language (Polish) using supervised fine-tuning. The dataset assigns speeches to binary ideological categories (Left-wing and Right-wing) based on the speakers'… See the full description on the dataset page: https://huggingface.co/datasets/PoliWings/Polish-Parliamentary-Speeches-Corpus.text10K<n<100K0 likes24 downloads4mo agoHugging Face03Decima-Data /luxembourgish-parliamentary-corpus Luxembourgish Parliamentary Corpus (2023–2028) A provenance-documented, speaker-attributed corpus of Luxembourg's parliamentary proceedings, built from the official session reports (comptes rendus / "D'Chamberblietchen") of the Chambre des Députés, legislature 2023–2028. Luxembourgish (Lëtzebuergesch) is a documented low-resource language: the Luxembourgish Wikipedia holds roughly 64,000 articles and most large language models perform poorly in it for lack of training material.… See the full description on the dataset page: https://huggingface.co/datasets/Decima-Data/luxembourgish-parliamentary-corpus.texttext-generation100K<n<1M0 likes23 downloads3mo agoHugging Face04balaji-ramk /parliamentary-debate-casestextn<1K0 likes5 downloads2y agoHugging Face05ascentftw /nation-parliamentary-promptstabular1K<n<10K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.