Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jkot /parliament_hearings_processed Preprocessed parliament hearings ASR dataset to truecased form. Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126 dataset_info: features: - name: id dtype: string - name: audio dtype: audio: sampling_rate: 16000 - name: transcription sequence: string splits: - name: train num_bytes: 53645064353.18 num_examples: 191455 - name: test num_bytes: 740331298.0 num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.audio100K<n<1M1 likes20k downloads3y agoHugging Face02rempah /malaysia-parliament-bills0 likes1.6k downloads1y agoHugging Face03NickWeng /tw_parliament_split Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.audioautomatic-speech-recognition100K<n<1M0 likes769 downloads20d agoHugging Face04manifoldix /swg_parliament_fhnwtext10K<n<100K2 likes727 downloads5y agoHugging Face05OpenFormosa /parliament Parliament Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the zh_tw split of disco-eth/WorldSpeech. It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD parliamentary proceedings. This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style text filtering pipeline. Audio is preserved from the upstream dataset and cast as a Hugging Face Audio(sampling_rate=24000) feature. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/parliament.audioautomatic-speech-recognition100K<n<1M2 likes693 downloads4mo agoHugging Face06irf23 /canadian-parliamentary-expenditures Canadian House of Commons Parliamentary Expenditures Dataset This dataset contains detailed expenditure records from the Canadian House of Commons, spanning from 2021 Q2 to 2025 Q4, with 1,219,648 total expenditure records across 450 parliament members. Dataset Structure parliamentary_data_hf/ ├── data/ │ ├── train/ # Training split (2021-2024) │ │ ├── expenditures-2021-q2.parquet │ │ ├── expenditures-2021-q3.parquet │ │ ├── ...… See the full description on the dataset page: https://huggingface.co/datasets/irf23/canadian-parliamentary-expenditures.tabulartabular-classification1M<n<10M0 likes671 downloads1y agoHugging Face07yanickschraner /swiss_parliament_corpus Dataset Card for "swiss_parliament_corpus" More Information needed audio10K<n<100K1 likes543 downloads4y agoHugging Face08boun-tabilab /turkish_parliamentary_data Grand National Assembly Corpus of Türkiye (GNACT) A comprehensive collection of Turkish parliamentary transcripts spanning over 100 years (1920–present), from 10 legislative bodies. Includes both Ottoman Turkish (1920–1928) and Modern Turkish (1928–present) texts. Loading the dataset from datasets import load_dataset # Strategy 1: full session documents, all bodies (default) ds = load_dataset("boun-tabilab/turkish_parliamentary_data", "full_sessions", split="train") #… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/turkish_parliamentary_data.tabulartext-generation1M<n<10M8 likes476 downloads6mo agoHugging Face09NbAiLab /norwegian_parliament Dataset Card Creation Guide Dataset Summary This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.texttext-classification1K<n<10K5 likes340 downloads2y agoHugging Face10malaysia-ai /malaysia-parliament-youtube Malaysia Parliament Youtube Entire videos from https://www.youtube.com/@PARLIMENMALAYSIA1 including Live. With total 2068 audio files, total 3985.4 hours. how to download huggingface-cli download --repo-type dataset \ --include '*.z*' \ --local-dir './' \ malaysia-ai/malaysia-parliament-youtube wget https://www.7-zip.org/a/7z2301-linux-x64.tar.xz tar -xf 7z2301-linux-x64.tar.xz ~/7zz x parlimen.zip -y -mmt40 Source code Source code at… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysia-parliament-youtube.1 likes311 downloads1y agoHugging Face11gttsehu /basque_parliament_1 Dataset Card for Basque Parliament Speech Corpus 1.0 This work was partially funded by the Spanish Ministry of Science and Innovation (OPENSPEECH project, PID2019-106424RB-I00). Dataset Summary The Basque Parliament Speech Corpus 1.0 consists of 1462 hours of speech extracted from Basque Parliament plenary sessions from 2013 to 2022. Encoded as MP3 files, the dataset contains 759192 transcribed segments either spoken in Basque, Spanish or both (in Basque and Spanish).… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/basque_parliament_1.automatic-speech-recognition3 likes285 downloads2y agoHugging Face12TomData /speeches-of-the-german-parliament1 likes247 downloads8mo agoHugging Face13parliament-data /hansard-embeddings Hansard Embeddings & Derived Analytics Text embeddings and derived analytics cubes covering the parliamentary Hansard of four Australian jurisdictions: the Commonwealth (au), New South Wales (nsw), South Australia (sa), and Western Australia (wa). No Hansard prose is in this dataset. It ships vectors, join keys, and derived metadata (subject headings, member names, parties, counts). The official record remains with each parliament; to hydrate the text, run the open-source… See the full description on the dataset page: https://huggingface.co/datasets/parliament-data/hansard-embeddings.text1M<n<10M0 likes203 downloads3mo agoHugging Face14mteb /norwegian_parliamenttext1K<n<10K0 likes165 downloads1y agoHugging Face15AndrewMcDowell /de_corpora_parliament_processedtext100K<n<1M2 likes158 downloads5y agoHugging Face16emeierkeio /parliamentrag-camera-leg19 ParliamentRAG — Italian Chamber of Deputies, 19th legislature Full proceedings of the Italian Chamber of Deputies (Camera dei deputati) for the 19th legislature, from the first sitting on 13 October 2022 through 6 August 2026: verbatim speech transcripts, roll-call votes with every individual ballot, parliamentary acts with EuroVoc subjects, and the deputies' group, committee and government memberships over time. The dataset is refreshed as new sittings are ingested. The tables… See the full description on the dataset page: https://huggingface.co/datasets/emeierkeio/parliamentrag-camera-leg19.tabulartext-retrieval1M<n<10M0 likes153 downloads1mo agoHugging Face17Iskaj /dutch_corpora_parliament_processedtext1M<n<10M1 likes144 downloads5y agoHugging Face18JonathanSum /en_corpora_parliament_processedtext1M<n<10M0 likes143 downloads5y agoHugging Face19Plim /fr_corpora_parliament_processedtext1M<n<10M0 likes142 downloads4y agoHugging Face20HarrisDePerceptron /sv_corpora_parliament_processedtext1M<n<10M0 likes141 downloads5y agoHugging Face21RuudVelo /nl_corpora_parliament_processedtext1M<n<10M1 likes141 downloads5y agoHugging Face22JonathanSum /sv_corpora_parliament_processedtext1M<n<10M0 likes137 downloads5y agoHugging Face23LuisG07 /es_corpora_parliament_processedtext1M<n<10M0 likes134 downloads5y agoHugging Face24anushakamath /sv_corpora_parliament_processed_v0text1M<n<10M0 likes125 downloads5y agoHugging Face25azuur /es_corpora_parliament_processedtext1M<n<10M0 likes124 downloads5y agoHugging Face26emilpartow /german-parliament-speeches German Parliament Speeches This dataset contains speeches from the German parliament, derived from the Open Discourse Project (Harvard Dataverse). Source Data source: Open Discourse ProjectHarvard DataverseDOI: 10.7910/DVN/FIKIBO Original citation: @data{DVN/FIKIBO_2020, author = {Richter, Florian and Koch, Philipp and Franke, Oliver and Kraus, Jakob and Kuruc, Fabrizio and Thiem, Anja and Högerl, Judith and Heine, Stella and Schöps, Konstantin}, publisher = {Harvard… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/german-parliament-speeches.tabulartext-classification100K<n<1M4 likes119 downloads1y agoHugging Face27DimitarV /bulgarian-parliament-moss-1to1 Bulgarian Parliament MOSS 1-to-1 Private machine-labelled training candidates under construction. Not gold data or a held-out evaluation set. Public redistribution terms remain unverified. Source: DimitarV/eurospeech-bg-single-speaker at 6aa43432891867bcd249c0fe944ac6286888231b, derived from disco-eth/EuroSpeech. Reference transcripts are parliamentary stenographic text. Speaker identities are inferred clusters, not named or human-verified speakers. Active batch list… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-parliament-moss-1to1.automatic-speech-recognition0 likes119 downloads26d agoHugging Face28JayJayThrowThrow /uk-parliament-hansard-modern uk-parliament-hansard-modern (FineWeb-style) Modern UK Hansard debates exported into LLM-training-friendly FineWeb-style Parquet shards. What’s inside Format: nanochat-parquet-v1 Layout: shard_*.parquet + metadata.json Text column: text Parquet settings: zstd (level 3), row_group_size=1024, use_dictionary=False, write_statistics=False HuggingFace Dataset repo: JayJayThrowThrow/uk-parliament-hansard-modern Loading from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/JayJayThrowThrow/uk-parliament-hansard-modern.text100K<n<1M0 likes114 downloads10mo agoHugging Face29Elormiden /Hellenic-greek-parliamentary-speech HParl: Hellenic Parliamentary Speech Corpus Dataset Description Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors. Link to the original source: https://inventory.clarin.gr/corpus/1602 HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.audio10K<n<100K1 likes100 downloads1y agoHugging Face30sl-parliamentary-nlp /HansardNER HansardNER HansardNER is a named entity recognition (NER) dataset for Sinhala parliamentary proceedings. It contains 4,000 speaker turns (1.49 million words) from 494 sitting days of the Sri Lankan Hansard, 2017–2026, labelled with nine entity classes. The dataset has two kinds of labels: Silver (all 4,000 turns): labelled by a large language model (Google Gemini) under written guidelines. Not checked by a person. Gold (648 of those turns): the silver labels corrected by human… See the full description on the dataset page: https://huggingface.co/datasets/sl-parliamentary-nlp/HansardNER.texttoken-classification1K<n<10K0 likes68 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.