Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M270 likes19k downloads3y agoHugging Face02google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M41 likes10k downloads3y agoHugging Face03tmquan /anle-toaan-gov-vn Vietnamese Án lệ Corpus — anle.toaan.gov.vn 🇻🇳 Tóm tắt. Bộ dữ liệu các bản án + án lệ Việt Nam thu thập từ cổng anle.toaan.gov.vn của Tòa án nhân dân tối cao. Mỗi văn bản đi kèm markdown chuẩn hoá tiếng Việt và một lớp grounding mức câu (mỗi trích dẫn mang sentence_id + char span trỏ ngược vào markdown). Bộ dữ liệu là một phần của ViLA common-corpus và ship ba cấu hình HF theo chuẩn chung: documents (bảng chính) · embeddings (vector 4096-D Nemotron-3-Embed-8B) · reduces (toạ… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/anle-toaan-gov-vn.tabulartext-classification10K<n<100K10 likes8k downloads24d agoHugging Face04IPEC-COMMUNITY /libero_goal_no_noops_1.0.0_lerobottabular10K<n<100K1 likes5.8k downloads11mo agoHugging Face05ZombitX64 /xauusd-gold-price-historical-data-2004-2025 XAUUSD Gold Price Historical Data 2004-2025 This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025. Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024" Content: The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns: Date Open High Low Close Volume Usage: This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.tabular1M<n<10M12 likes2.8k downloads1y agoHugging Face06AgentPublic /open_government Open Government Dataset Open Government is the largest agregation of governement text and data made available as part of open data programs. In total, the dataset contains approximately 380B tokens. While Open Government aims to become a global resource, in its current state it mostly features open datasets from the US, France, European and international organizations. The dataset comprises 16 collections curated through two different initiaties: Finance commons and Legal commons.… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/open_government.tabulartext-generation10M<n<100M5 likes2.4k downloads2y agoHugging Face07pfaha /goodreads-books Goodreads Books Metadata Dataset Description Goodreads Books Metadata is a structured dataset of book records scraped directly from Goodreads, a social platform for book readers and recommendations. The dataset is collected via an ongoing, resumable crawl and contains rich metadata per book: bibliographic information, crowd-sourced ratings, contributor (author/illustrator/editor/etc.) details enriched with author-level popularity stats, genre tags, series… See the full description on the dataset page: https://huggingface.co/datasets/pfaha/goodreads-books.tabulartabular-regression100K<n<1M1 likes2.1k downloads23h agoHugging Face08GonzaloA /fake_news TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: https://huggingface.co/spaces/huggingface/datasets-tagging annotations_creators: - no-annotation language_creators: - found language: - en license: - unknown multilinguality: - monolingual size_categories: - 30k<n<50k source_datasets: - original task_categories: - text-classification task_ids: - fact-checking - intent-classification pretty_name: GonzaloA / Fake News Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/GonzaloA/fake_news.tabular10K<n<100K28 likes2k downloads4y agoHugging Face09BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes1.9k downloads10mo agoHugging Face10vidore /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K1 likes1.6k downloads1y agoHugging Face11ZipLime /us-government-contracts US Federal Contract Awards Federal contract actions paid to publicly traded companies — and, for each one, the day the public could first see it, which for the Department of Defense is three months after the contract was signed. 13 270 259 contract transactions · 10 445 332 awards · 1 700 listed companies · FY2005 to today · $4.42tn obligated 69% of those rows are the Department of Defense and the Army Corps of Engineers, and not one of them is dated less than ninety-two days… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/us-government-contracts.tabulartabular-regression10M<n<100M0 likes1.4k downloads13h agoHugging Face12google /MusicCaps Dataset Card for MusicCaps Dataset Summary The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g., "A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.tabulartext-to-speech1K<n<10K154 likes1.3k downloads4y agoHugging Face13gridfm /opf_small_case2000_gocThis is a dataset generated with gridfm-datakit for the 2000-bus system with the config, using the repro branch of datakit. Data download through hfApi Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times. tabulartabular-regression1B<n<10B0 likes1.3k downloads9d agoHugging Face14tmquan /nso-gov-vn nso-gov-vn — Vietnam NSO PX-Web mirror · Bản sao PX-Web của Tổng cục Thống kê 🇻🇳 Tóm tắt. Bản sao đầy đủ, từng bảng một, của cơ sở dữ liệu thống kê PX-Web của Tổng cục Thống kê Việt Nam (NSO / GSO) tại https://pxweb.nso.gov.vn. Mỗi ma trận PX-Web (multi-dimensional data cube) được phơi ra cùng lúc ở (i) bản gốc với schema riêng và (ii) định dạng long-format gộp chung, để bạn có thể chọn giữa "nguyên trạng" hay "join sẵn". 🇬🇧 Summary. A complete, table-by-table mirror of the… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/nso-gov-vn.tabulartabular-classification100K<n<1M0 likes1.3k downloads5mo agoHugging Face15gridfm /pf_small_case10000_gocThis is a dataset generated with gridfm-datakit for the 10000-bus system with the config, using the repro branch of datakit. tabular10B<n<100B0 likes1.3k downloads9d agoHugging Face16google /code_x_glue_cc_clone_detection_big_clone_bench Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench" Dataset Summary CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score. The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.tabulartext-classification1M<n<10M22 likes1.2k downloads3y agoHugging Face17google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.1k downloads3y agoHugging Face18Godners /DaoZang DaoZang — 道藏经文语料数据集 《中华道藏》《正统道藏》整理本语料的可加载数据集,与向量库 (ChromaDB) 逐块对应, 供检索评测、微调与 RAG 使用。 数据分片 分片 行数 粒度 字段 train (data/train-00000-of-00006.parquet ~ ...00005-of-00006.parquet, 6 个分片) 285,117 文本块 (chunk) source / title / chunk_index / chars / text / embedding 标准分片命名 (train-XXXXX-of-00006.parquet),load_dataset 自动合并,无需改动; 每片约 217MB(含 20,000 行一个 row group 与 page index),方便 HF Dataset Viewer 在线浏览(单次扫描上限 ~300MB); source = 源 Markdown 文件名 (与 ChromaDB 元数据一致);… See the full description on the dataset page: https://huggingface.co/datasets/Godners/DaoZang.tabulartext-retrieval100K<n<1M0 likes984 downloads19d agoHugging Face19tmquan /phapdien-moj-gov-vn Bộ Pháp Điển Việt Nam — phapdien.moj.gov.vn 🇻🇳 Tóm tắt. Bộ ngữ liệu cấp Điều của Bộ Pháp Điển Việt Nam — bộ pháp điển chính thức do Bộ Tư pháp công bố. Mỗi dòng documents là một Điều kèm toàn văn đã chuẩn hoá, chương sở thuộc, đề mục và chủ đề. Kèm theo là vector nhúng ngữ nghĩa 4096-D (embeddings), toạ độ giảm chiều trong không gian chung ViLA (reduces), và từ điển ontology song ngữ Việt–Anh (chủ đề · đề mục · thuật ngữ). 🇬🇧 One-line. Article-level corpus of the Bộ Pháp… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/phapdien-moj-gov-vn.imagetext-classification100K<n<1M12 likes963 downloads24d agoHugging Face20liyucheng /goodreadstabular100M<n<1B3 likes947 downloads2y agoHugging Face21csoai /gspc-gov GSPC — governance bank (GovBench) In one line: The frozen EU AI Act risk-tier classification bank (GovBench) behind the board's governance axis. For evaluators who want to test a model on the same items. Use it from datasets import load_dataset ds = load_dataset("csoai/gspc-gov", split="train") print(ds[0]) Verify a signed card in your browser, free, no account: https://councilof.ai/gspc-verify/?ref=hf-gspc-gov For agents: MCP endpoint POST… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-gov.tabularquestion-answeringn<1K0 likes923 downloads7h agoHugging Face22gridfm /pf_small_case2000_gocThis is a dataset generated with gridfm-datakit for the 2000-bus system with the config, using the repro branch of datakit. Data download through hfApi Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times. tabulartabular-regression1B<n<10B0 likes862 downloads9d agoHugging Face23gridfm /opf_small_case10000_gocThis is a dataset generated with gridfm-datakit for the 10000-bus system with the config, using the repro branch of datakit. tabular100M<n<1B0 likes826 downloads9d agoHugging Face24csoai /gspc-jail-goldbank GSPC — jail bank (GoldBank-Detector) Council of AI measurement bank. Measurement, not certification. Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Live measurement. This bank stands behind the jail row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=jail (family, kind, status and n are on that row, never typed here; the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-jail-goldbank.tabularquestion-answeringn<1K0 likes798 downloads5d agoHugging Face25mteb /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K0 likes788 downloads8mo agoHugging Face26sfc-gh-goliaro /wildchat-mixed-1k wildchat-mixed-1k Real-world chat requests for end-to-end LLM inference benchmarking in fastkernels — Scenario A, the bulk-throughput workload used to saturate continuous batching with a realistic mix of short/long prompts and short/long responses. What it's for One dataset that replaces separate prefill-heavy / balanced / decode-heavy splits: its natural length distribution puts prefill-bound and decode-bound requests in the same batch, so a single run yields a… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/wildchat-mixed-1k.tabulartext-generation1K<n<10K0 likes786 downloads3mo agoHugging Face27kl41r3 /erc8004-vs-a2a-governance Agentic Protocol Governance · v2.0.0 The current, self-contained release is releases/v2.0.0/. It supplies the public inputs, fixed annotations, reference results, locked analysis code and one reproduction entry point for the submitted study. You do not need another Git repository or the private research workspace. Download and reproduce Requirements: Python 3.14, uv, and make. Dependency installation may use the network; the analyses make no hosted-model calls.… See the full description on the dataset page: https://huggingface.co/datasets/kl41r3/erc8004-vs-a2a-governance.tabular10K<n<100K0 likes781 downloads1mo agoHugging Face28orailix /ride-gold-standard RIDE Gold Standard RIDE Gold Standard is the full benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations. This release is intended as the primary benchmark tier for RIDE. It is used for full-scale evaluation and comparison of models under the shared RIDE prediction task and evaluation protocol. Links… See the full description on the dataset page: https://huggingface.co/datasets/orailix/ride-gold-standard.tabular1M<n<10M0 likes774 downloads4mo agoHugging Face29goodcoffee /Meter_Readingimage1K<n<10K3 likes744 downloads2y agoHugging Face30RIPS-Goog-23 /IIT-CDIP Dataset Card for "IIT-CDIP-2" More Information needed tabular1M<n<10M10 likes723 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.