datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GraphRAG-Bench
GraphRAG-Bench : A Comprehensive Benchmark for Evaluating Graph Retrieval-Augmented Generation Models
🎉News •
📖About •
🏆Leaderboards •
🧩Task Examples
🔧Getting Started •
📬Contact •
📝Citation
This repository is for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG… See the full description on the dataset page: https://huggingface.co/datasets/GraphRAG-Bench/GraphRAG-Bench.szl-estate-graph
SZL Estate Graph v1
This is a deterministic publication bundle for SZLHOLDINGS/szl-estate-graph.
It turns the two receipted SZL Constellation topology documents into a typed,
trinity-connected graph suitable for graph-learning and drift comparison.
Publication target: SZLHOLDINGS/szl-estate-graph. Hub publication and its
immutable revision are provider evidence separate from this content bundle;
neither publication nor download implies model training, admission, or runtime
use.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-estate-graph.MUTAG
Dataset Card for MUTAG
Dataset Summary
The MUTAG dataset is 'a collection of nitroaromatic compounds and the goal is to predict their mutagenicity on Salmonella typhimurium'.
Supported Tasks and Leaderboards
MUTAG should be used for molecular property prediction (aiming to predict whether molecules have a mutagenic effect on a given bacterium or not), a binary classification task. The score used is accuracy, using a 10-fold cross-validation.
External… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/MUTAG.sketch-graph-vitruvion
pretty_name: SketchGraphs, Vitruvion selection (rebuilt)
license: other
license_name: onshape-terms-of-use
license_link: https://www.onshape.com/legal/terms-of-use#your_content
size_categories:
- 1M<n<10M
tags:
- cad
- parametric-cad
- sketches
- geometric-constraints
- sketchgraphs
- vitruvion
SketchGraphs, Vitruvion selection (sg_filtered_unique.npy, rebuilt)
This is the dataset of Vitruvion (Seff et al., Vitruvion: A Generative Model of… See the full description on the dataset page: https://huggingface.co/datasets/benikm91/sketch-graph-vitruvion.GraphQAtrajectory-graph-monitoring
Trajectory Graph Monitoring v0.2; Research Release
Prospective boundary prediction from cumulative typed trajectory transitionsR.J. Sabouhi · Symbolic Suite · August 2026
What this is
Trajectory Graph Monitoring is a prospective runtime-evaluation benchmark. It asks whether structural information available in an incomplete execution trajectory can improve prediction of a boundary violation that occurs later. Every evaluated prefix ends before the violating action… See the full description on the dataset page: https://huggingface.co/datasets/rjsabouhi/trajectory-graph-monitoring.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.Nemotron-Problem-Graph-v2federal-register-live-graphrag-research-20260810
Federal Register live GraphRAG (research)
Local LCR-071 live pipeline output for the 2026-08-10 cutoff (11,784 documents,
CUDA thenlper/gte-small). This Hub copy is a research snapshot.
It is not a current-bundle and does not replace
justicedao/ipfs_federal_register. LCR-084 remains open. Official Federal
Register publications remain the authority.
Hub git directories may contain at most 10,000 files. Document bodies beyond
that cap are stored under corpus/bodies-part2/ rather… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-live-graphrag-research-20260810.MTID
TurnGate: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Overview
TurnGate is a response-aware defense mechanism designed to detect and mitigate hidden malicious intent in multi-turn dialogue systems. Defending state-of-the-art multi-turn malicious attacks like CKA-Agent.
MTID Dataset
We include the MTID (Multi-Turn Intent Dataset) here. This dataset contains a collection of… See the full description on the dataset page: https://huggingface.co/datasets/Graph-COM/MTID.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_knowledge_graph_qa.bottleneck-oracle-graphsPROTEINS
Dataset Card for PROTEINS
Dataset Summary
The PROTEINS dataset is a medium molecular property prediction dataset.
Supported Tasks and Leaderboards
PROTEINS should be used for molecular property prediction (aiming to predict whether molecules are enzymes or not), a binary classification task. The score used is accuracy, using a 10-fold cross-validation.
External Use
PyGeometric
To load in PyGeometric, do the following:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/PROTEINS.GraphRAG
Local GraphRAG Research Artifacts
This repository contains the data and output artifacts produced for my master's
thesis research on knowledge graph construction in local GraphRAG systems. It
is an evidence package for the study, rather than the code used to run the
experiment.
The collection includes the study corpus, model extraction outputs, final
knowledge graphs, retrieved evidence, generated answers, and evaluation results.
Together, these files make it possible to inspect… See the full description on the dataset page: https://huggingface.co/datasets/boblaros/GraphRAG.GraphInstruct-Testawesome-graph-engineering
Awesome Graph Engineering Resource Atlas
A versioned collection of research, standards, frameworks, protocols, reliability systems, evaluations, and critiques for graph-structured multi-agent systems and programmable AI-agent organizations.
This dataset mirrors Awesome Graph Engineering. The GitHub JSONL file is canonical; the Hub exposes the same records through Dataset Viewer, direct downloads, datasets, and pandas.
Working definition
Graph engineering is the… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-graph-engineering.Scientific-Citation-Graph
RegalFire Scientific reference graph
RegalFire — AI Data Foundry
Literal external URLs and DOI-like references with mechanically extracted text context and complete forum attribution.
Verified scope
Records: 888; distinct source threads: 221; unique answers represented: 320.
Domain thread counts: {"statistics": 61, "computational_science": 71, "biology": 89}.
Actual record splits: {"train": 543, "test": 109, "validation": 109, "holdout": 127}.
Each record includes… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Scientific-Citation-Graph.sketch-graph-xl
Primitive
type
coordinates_params
Canonical order
Line
PartLineWithId
[start, end]
Left point first; if the x values are within 0.01, the upper point first
Circle
Circle
[leftmost, rightmost]
Left, then right; both have the centre's y
Arc
Arc
[start, mid, end]
Runs clockwise on the image from start to end; mid is the point halfway along the arc
(end marker)
FinishDrawing
none
Always last
qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.IMDB-BINARY
Dataset Card for IMDB-BINARY (IMDb-B)
Dataset Summary
The IMDb-B dataset is "a movie collaboration dataset that consists of the ego-networks of 1,000 actors/actresses who played roles in movies in IMDB. In each graph, nodes represent actors/actress, and there is an edge between them if they appear in the same movie. These graphs are derived from the Action and Romance genres".
Supported Tasks and Leaderboards
IMDb-B should be used for graph classification… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/IMDB-BINARY.graph-captioning-train-onlylex-au-graph
lex-au-graph
Cross-reference knowledge graph over Australian Commonwealth Acts, built from
the lex-au AKN 3.0 XML corpus
— the retrieval layer of the AU Legislative Intelligence Stack.
This dataset publishes three files, rebuilt automatically whenever lex-au
publishes a corpus update:
graph.json — the full cross-reference graph (nodes: Acts, Sections,
defined terms; edges: contains, defines, ref, mentions), as NetworkX
node_link_data JSON. Not tabular — load it directly, not… See the full description on the dataset page: https://huggingface.co/datasets/cchew/lex-au-graph.CATH4.2fin-cfa-graphgen
Fin-CFA-GraphGen: 785K Knowledge-Guided Financial QA Examples
Fin-CFA-GraphGen is a large-scale English dataset for financial instruction tuning, financial question answering, and domain-specific language-model post-training. It contains 785,149 synthetic question–answer examples generated from CFA curriculum and exam-preparation books with the GraphGen knowledge-driven data-generation method.
The dataset and its role in the post-training pipeline are described in Data-Centric… See the full description on the dataset page: https://huggingface.co/datasets/whoisjiji/fin-cfa-graphgen.Graph2Counsel
Dataset Card
Paper: Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs
Language(s) (NLP): English
license: cc-by-sa-4.0
Dataset Summary
Graph2Counsel is a synthetic counseling session dataset generated from Client Psychological Graphs (CPGs). The dataset provides the CPG, generated diverse client profiles and dialogues from this CPG as well as counselor strategies corresponding to each CPG collected from the real… See the full description on the dataset page: https://huggingface.co/datasets/UKPLab/Graph2Counsel.sketch-graph
GraphMastermedimate-ddi-graph-source-data
MediMate DDInter graph source data
Patient-free, source-reported positive DDInter pairs for a conditional three-grade task. train.jsonl has 48,212 pairs; dev.jsonl has 1,720; holdout.jsonl has 1,633. Drug identities and parent structures do not overlap across the splits. drugs.jsonl has 1,532 source-verified molecular structures. Unknown or unlisted pairs are not negative examples.
source_atc_v1.json contains source-page ATC codes. pharmacology.jsonl contains openFDA… See the full description on the dataset page: https://huggingface.co/datasets/nthan2005/medimate-ddi-graph-source-data.PKU-Alignment-Graphdouvras-scientific-ci-evidence-graph
Douvras Scientific CI Evidence Graph v0.1
Synthetic protocol dataset for linking a claim to its paper, repository,
dataset, seed and reproduced metric. It contains 30 records from six toy paper
instances (20 train, 5 validation and 5 frozen test), split by paper_id.
The labels distinguish REPRODUCED, PARTIAL, FAILED and INCONCLUSIVE.
Shortcuts and leakage fail closed. No real paper, code, dataset or result is
included, and this release is not a reproduction benchmark.
