datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.civil_comments
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.anle-toaan-gov-vn
Vietnamese Án lệ Corpus — anle.toaan.gov.vn
🇻🇳 Tóm tắt. Bộ dữ liệu các bản án + án lệ Việt Nam thu thập từ cổng
anle.toaan.gov.vn của Tòa án nhân dân tối cao.
Mỗi văn bản đi kèm markdown chuẩn hoá tiếng Việt và một lớp grounding mức câu
(mỗi trích dẫn mang sentence_id + char span trỏ ngược vào markdown). Bộ dữ
liệu là một phần của ViLA common-corpus và ship ba cấu hình HF theo chuẩn
chung: documents (bảng chính) · embeddings (vector 4096-D Nemotron-3-Embed-8B) ·
reduces (toạ… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/anle-toaan-gov-vn.libero_goal_no_noops_1.0.0_lerobotxauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.open_government
Open Government Dataset
Open Government is the largest agregation of governement text and data made available as part of open data programs.
In total, the dataset contains approximately 380B tokens. While Open Government aims to become a global resource, in its current state it mostly features open datasets from the US, France, European and international organizations.
The dataset comprises 16 collections curated through two different initiaties: Finance commons and Legal commons.… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/open_government.goodreads-books
Goodreads Books Metadata
Dataset Description
Goodreads Books Metadata is a structured dataset of book records scraped directly from Goodreads, a social platform for book readers and recommendations.
The dataset is collected via an ongoing, resumable crawl and contains rich metadata per book: bibliographic information, crowd-sourced ratings, contributor (author/illustrator/editor/etc.) details enriched with author-level popularity stats, genre tags, series… See the full description on the dataset page: https://huggingface.co/datasets/pfaha/goodreads-books.fake_news
TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: https://huggingface.co/spaces/huggingface/datasets-tagging
annotations_creators:
- no-annotation
language_creators:
- found
language:
- en
license:
- unknown
multilinguality:
- monolingual
size_categories:
- 30k<n<50k
source_datasets:
- original
task_categories:
- text-classification
task_ids:
- fact-checking
- intent-classification
pretty_name: GonzaloA / Fake News
Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/GonzaloA/fake_news.govdocs1-pdf-source
govdocs1: source PDF files
[!NOTE]
Converted versions of other document types (word, txt, etc) are available in this repo
This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd.
Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details
5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
us-government-contracts
US Federal Contract Awards
Federal contract actions paid to publicly traded companies — and, for each one,
the day the public could first see it, which for the Department of Defense
is three months after the contract was signed.
13 270 259 contract transactions · 10 445 332 awards · 1 700 listed companies ·
FY2005 to today · $4.42tn obligated
69% of those rows are the Department of Defense and the Army Corps of
Engineers, and not one of them is dated less than ninety-two days… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/us-government-contracts.MusicCaps
Dataset Card for MusicCaps
Dataset Summary
The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g.,
"A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.opf_small_case2000_gocThis is a dataset generated with gridfm-datakit for the 2000-bus system with the config, using the repro branch of datakit.
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
nso-gov-vn
nso-gov-vn — Vietnam NSO PX-Web mirror · Bản sao PX-Web của Tổng cục Thống kê
🇻🇳 Tóm tắt. Bản sao đầy đủ, từng bảng một, của cơ sở dữ liệu thống
kê PX-Web của Tổng cục Thống kê Việt Nam (NSO / GSO) tại
https://pxweb.nso.gov.vn. Mỗi ma trận PX-Web (multi-dimensional
data cube) được phơi ra cùng lúc ở (i) bản gốc với schema riêng và
(ii) định dạng long-format gộp chung, để bạn có thể chọn giữa
"nguyên trạng" hay "join sẵn".
🇬🇧 Summary. A complete, table-by-table mirror of the… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/nso-gov-vn.pf_small_case10000_gocThis is a dataset generated with gridfm-datakit for the 10000-bus system with the config, using the repro branch of datakit.
code_x_glue_cc_clone_detection_big_clone_bench
Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench"
Dataset Summary
CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench
Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score.
The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.discofuse
Dataset Card for "discofuse"
Dataset Summary
DiscoFuse is a large scale dataset for discourse-based sentence fusion.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
discofuse-sport
Size of downloaded dataset files: 4.33 GB
Size of the generated dataset: 15.04 GB
Total amount of disk used: 19.36 GB
An example of 'train' looks as follows.
{… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.DaoZang
DaoZang — 道藏经文语料数据集
《中华道藏》《正统道藏》整理本语料的可加载数据集,与向量库 (ChromaDB) 逐块对应,
供检索评测、微调与 RAG 使用。
数据分片
分片
行数
粒度
字段
train (data/train-00000-of-00006.parquet ~ ...00005-of-00006.parquet, 6 个分片)
285,117
文本块 (chunk)
source / title / chunk_index / chars / text / embedding
标准分片命名 (train-XXXXX-of-00006.parquet),load_dataset 自动合并,无需改动;
每片约 217MB(含 20,000 行一个 row group 与 page index),方便 HF Dataset Viewer
在线浏览(单次扫描上限 ~300MB);
source = 源 Markdown 文件名 (与 ChromaDB 元数据一致);… See the full description on the dataset page: https://huggingface.co/datasets/Godners/DaoZang.phapdien-moj-gov-vn
Bộ Pháp Điển Việt Nam — phapdien.moj.gov.vn
🇻🇳 Tóm tắt. Bộ ngữ liệu cấp Điều của Bộ Pháp Điển Việt Nam — bộ pháp điển
chính thức do Bộ Tư pháp công bố. Mỗi dòng documents là một Điều kèm toàn văn đã
chuẩn hoá, chương sở thuộc, đề mục và chủ đề. Kèm theo là vector nhúng ngữ nghĩa 4096-D
(embeddings), toạ độ giảm chiều trong không gian chung ViLA (reduces), và từ điển
ontology song ngữ Việt–Anh (chủ đề · đề mục · thuật ngữ).
🇬🇧 One-line. Article-level corpus of the Bộ Pháp… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/phapdien-moj-gov-vn.goodreadsgspc-gov
GSPC — governance bank (GovBench)
In one line: The frozen EU AI Act risk-tier classification bank (GovBench) behind the board's governance axis. For evaluators who want to test a model on the same items.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-gov", split="train")
print(ds[0])
Verify a signed card in your browser, free, no account: https://councilof.ai/gspc-verify/?ref=hf-gspc-gov
For agents: MCP endpoint POST… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-gov.pf_small_case2000_gocThis is a dataset generated with gridfm-datakit for the 2000-bus system with the config, using the repro branch of datakit.
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opf_small_case10000_gocThis is a dataset generated with gridfm-datakit for the 10000-bus system with the config, using the repro branch of datakit.
gspc-jail-goldbank
GSPC — jail bank (GoldBank-Detector)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the jail row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=jail (family, kind, status and n are on that row, never typed here; the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-jail-goldbank.syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
wildchat-mixed-1k
wildchat-mixed-1k
Real-world chat requests for end-to-end LLM inference benchmarking in fastkernels — Scenario A, the bulk-throughput workload used to saturate continuous batching with a realistic mix of short/long prompts and short/long responses.
What it's for
One dataset that replaces separate prefill-heavy / balanced / decode-heavy splits: its natural length distribution puts prefill-bound and decode-bound requests in the same batch, so a single run yields a… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/wildchat-mixed-1k.erc8004-vs-a2a-governance
Agentic Protocol Governance · v2.0.0
The current, self-contained release is releases/v2.0.0/. It supplies the public inputs, fixed annotations, reference results, locked analysis code and one reproduction entry point for the submitted study. You do not need another Git repository or the private research workspace.
Download and reproduce
Requirements: Python 3.14, uv, and make. Dependency installation may use the network; the analyses make no hosted-model calls.… See the full description on the dataset page: https://huggingface.co/datasets/kl41r3/erc8004-vs-a2a-governance.ride-gold-standard
RIDE Gold Standard
RIDE Gold Standard is the full benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations.
This release is intended as the primary benchmark tier for RIDE. It is used for full-scale evaluation and comparison of models under the shared RIDE prediction task and evaluation protocol.
Links… See the full description on the dataset page: https://huggingface.co/datasets/orailix/ride-gold-standard.Meter_ReadingIIT-CDIP
Dataset Card for "IIT-CDIP-2"
More Information needed
