datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trajectory-graph-monitoring
Trajectory Graph Monitoring v0.2; Research Release
Prospective boundary prediction from cumulative typed trajectory transitionsR.J. Sabouhi · Symbolic Suite · August 2026
What this is
Trajectory Graph Monitoring is a prospective runtime-evaluation benchmark. It asks whether structural information available in an incomplete execution trajectory can improve prediction of a boundary violation that occurs later. Every evaluated prefix ends before the violating action… See the full description on the dataset page: https://huggingface.co/datasets/rjsabouhi/trajectory-graph-monitoring.federal-register-live-graphrag-research-20260810
Federal Register live GraphRAG (research)
Local LCR-071 live pipeline output for the 2026-08-10 cutoff (11,784 documents,
CUDA thenlper/gte-small). This Hub copy is a research snapshot.
It is not a current-bundle and does not replace
justicedao/ipfs_federal_register. LCR-084 remains open. Official Federal
Register publications remain the authority.
Hub git directories may contain at most 10,000 files. Document bodies beyond
that cap are stored under corpus/bodies-part2/ rather… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-live-graphrag-research-20260810.MTID
TurnGate: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Overview
TurnGate is a response-aware defense mechanism designed to detect and mitigate hidden malicious intent in multi-turn dialogue systems. Defending state-of-the-art multi-turn malicious attacks like CKA-Agent.
MTID Dataset
We include the MTID (Multi-Turn Intent Dataset) here. This dataset contains a collection of… See the full description on the dataset page: https://huggingface.co/datasets/Graph-COM/MTID.bottleneck-oracle-graphsGraphRAG
Local GraphRAG Research Artifacts
This repository contains the data and output artifacts produced for my master's
thesis research on knowledge graph construction in local GraphRAG systems. It
is an evidence package for the study, rather than the code used to run the
experiment.
The collection includes the study corpus, model extraction outputs, final
knowledge graphs, retrieved evidence, generated answers, and evaluation results.
Together, these files make it possible to inspect… See the full description on the dataset page: https://huggingface.co/datasets/boblaros/GraphRAG.awesome-graph-engineering
Awesome Graph Engineering Resource Atlas
A versioned collection of research, standards, frameworks, protocols, reliability systems, evaluations, and critiques for graph-structured multi-agent systems and programmable AI-agent organizations.
This dataset mirrors Awesome Graph Engineering. The GitHub JSONL file is canonical; the Hub exposes the same records through Dataset Viewer, direct downloads, datasets, and pandas.
Working definition
Graph engineering is the… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-graph-engineering.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.lex-au-graph
lex-au-graph
Cross-reference knowledge graph over Australian Commonwealth Acts, built from
the lex-au AKN 3.0 XML corpus
— the retrieval layer of the AU Legislative Intelligence Stack.
This dataset publishes three files, rebuilt automatically whenever lex-au
publishes a corpus update:
graph.json — the full cross-reference graph (nodes: Acts, Sections,
defined terms; edges: contains, defines, ref, mentions), as NetworkX
node_link_data JSON. Not tabular — load it directly, not… See the full description on the dataset page: https://huggingface.co/datasets/cchew/lex-au-graph.fin-cfa-graphgen
Fin-CFA-GraphGen: 785K Knowledge-Guided Financial QA Examples
Fin-CFA-GraphGen is a large-scale English dataset for financial instruction tuning, financial question answering, and domain-specific language-model post-training. It contains 785,149 synthetic question–answer examples generated from CFA curriculum and exam-preparation books with the GraphGen knowledge-driven data-generation method.
The dataset and its role in the post-training pipeline are described in Data-Centric… See the full description on the dataset page: https://huggingface.co/datasets/whoisjiji/fin-cfa-graphgen.Graph2Counsel
Dataset Card
Paper: Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs
Language(s) (NLP): English
license: cc-by-sa-4.0
Dataset Summary
Graph2Counsel is a synthetic counseling session dataset generated from Client Psychological Graphs (CPGs). The dataset provides the CPG, generated diverse client profiles and dialogues from this CPG as well as counselor strategies corresponding to each CPG collected from the real… See the full description on the dataset page: https://huggingface.co/datasets/UKPLab/Graph2Counsel.GraphMasterPKU-Alignment-Graphdouvras-scientific-ci-evidence-graph
Douvras Scientific CI Evidence Graph v0.1
Synthetic protocol dataset for linking a claim to its paper, repository,
dataset, seed and reproduced metric. It contains 30 records from six toy paper
instances (20 train, 5 validation and 5 frozen test), split by paper_id.
The labels distinguish REPRODUCED, PARTIAL, FAILED and INCONCLUSIVE.
Shortcuts and leakage fail closed. No real paper, code, dataset or result is
included, and this release is not a reproduction benchmark.
fawkes-training-graph-embedded-260615
Fawkes — WM@Booth Training Graphs (v16 dataset) — PRIVATE
PRIVATE — derived from MIMIC-IV (PhysioNet credentialed, governed by the PhysioNet DUA). Do not redistribute. Credentialed access only.
The exact dataset the WM@Booth Graph-JEPA v16 model (on1onmangoes/fawkes-wmatbooth-graph-jepa-v16-260615) was trained on — 4,000 per-admission clinical knowledge graphs (~3,018 patients).
Each record = one hospital admission
field
what it is
subject_id, hadm_id… See the full description on the dataset page: https://huggingface.co/datasets/wmatbooth/fawkes-training-graph-embedded-260615.qualora-workforce-skills-graph
Qualora Workforce Skills Graph (Representative Sample)
Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/VegasIGN/qualora-workforce-skills-graph.praxis-chess-evidence-graphs
Praxis Chess Evidence Graphs
Built for, and used to
train and test,
praxis-chess-reasoner-qwen3.5-2b-lora
and its 4B sibling.
Chess mistakes from public Lichess data, each with an evidence graph of
engine-computed facts about the move, and a verified explanation: a typed
reasoning chain that a program can check claim by claim, the labels (what
happened, and why) and two or three sentences of prose.
The dataset exists to test one question: does a small model explain mistakes… See the full description on the dataset page: https://huggingface.co/datasets/praxis-chess/praxis-chess-evidence-graphs.Z3-Verified-Reasoning-Graphs
Z3-Verified Constraint Reasoning Dataset
5k Baseline · Production-Ready · Zero Label Noise
The Problem This Solves
Most synthetic reasoning datasets only show the "happy path". Real reasoning requires knowing when to backtrack.
Open-source LLMs hallucinate on constraint satisfaction problems because they are trained on fluent-sounding but logically inconsistent traces. This dataset is different:
❌ No LLM-generated reasoning — zero hallucinations, zero label noise
✅… See the full description on the dataset page: https://huggingface.co/datasets/nagygabor/Z3-Verified-Reasoning-Graphs.graph-reasoning-foundational-curriculum
Foundational Structural Curriculum — v1.1
TEMPORARY MIRROR. This repository is a convenience copy for browsing and
exploration, published so a reader can look at the data without S3/DVC
credentials. It is not the canonical dataset and is not guaranteed to
stay in sync. The source of truth is the DVC-tracked corpus in the
graph-reasoning-llm repo (pin by the git SHA of dvc.lock, never "latest").
The usual byte-for-byte verification against dvc.lock was skipped for
this upload.… See the full description on the dataset page: https://huggingface.co/datasets/akumch/graph-reasoning-foundational-curriculum.graphcontainer-graphs
GraphContainer Graph Artifacts
Overview
This repository contains preconstructed graph artifacts released with GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods.
Links:
Paper: https://arxiv.org/abs/2607.19362
Hugging Face paper page: https://huggingface.co/papers/2607.19362
Code: https://github.com/asmath472/GraphContainer
YouTube demo: https://youtu.be/O02eNJLwkU0
The Hugging Face datasets interface loads the artifact catalog in… See the full description on the dataset page: https://huggingface.co/datasets/hchaejeong/graphcontainer-graphs.phase3-graphs
Phase 3 Graphs
Graphs used in Phase 3
Data Card
Field
Type
Description
graph_id
string
Unique identifier for the graph
adjacency_matrix
List[List[int]]
NxN binary adjacency matrix where A[i,j]=1 means i→j
num_nodes
int
Number of nodes in the DAG
num_edges
int
Number of edges in the DAG
density
float
Graph density (edges / max_possible_edges)
method
string
Graph generation method (e.g., "PA", "ER")
Graph Generation Settings
Generation… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/phase3-graphs.clak-consulting-graph-ml
CLAK Consulting Knowledge Graph Dataset
Graph-ML dataset for consulting knowledge representation learning.
Dataset Structure
Inspired by OGB (Open Graph Benchmark) format:
Field
Type
Description
node_feat
list[list[int]]
Node features (type, one-hot)
edge_index
list[tuple[int,int]]
Edge pairs (source, target)
edge_attr
list[int]
Edge types
y
list[float]
Target labels
num_nodes
int
Node count
domain
str
Consulting domain
Node Types… See the full description on the dataset page: https://huggingface.co/datasets/Kraft102/clak-consulting-graph-ml.GraphWalkerBenchThis repository contains the GraphWalkerBench dataset from the paper GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum.
Code: https://github.com/XuShuwenn/GraphWalker
kgp-entity-citation-graph
King Pawn USA Entity and Citation Graph
Public-release-ready dataset card for operator-approved publication.
Dataset Summary
First-party King Pawn USA / King Gold and Pawn location facts, entity edges, citation targets, and citation evidence scan rows.
Data Fields
See dataset-metadata.json for file schemas.
Source Data
Location data comes from the canonical local KGP registry. Citation rows are reachable first-party URLs or reachable… See the full description on the dataset page: https://huggingface.co/datasets/CollateralAnalytics/kgp-entity-citation-graph.phase1-graphs
Phase 1 Graphs
Graphs used in Phase 1
Data Card
Field
Type
Description
graph_id
string
Unique identifier for the graph
adjacency_matrix
List[List[int]]
NxN binary adjacency matrix where A[i,j]=1 means i→j
num_nodes
int
Number of nodes in the DAG
num_edges
int
Number of edges in the DAG
density
float
Graph density (edges / max_possible_edges)
method
string
Graph generation method (e.g., "PA", "ER")
Graph Generation Settings
Generation… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/phase1-graphs.ai-collab-graph-2026
Ai Collab Graph 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-collab-graph-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-collab-graph-2026.GraphForge-SFT-2169-clean
GraphForge-SFT-2169
This release contains 2,169 trajectory sequences corresponding to the main
GraphForge SFT corpus for the 27B and 35B-A3B checkpoints. GLM-5.2 teacher
trajectories were admitted with agent-judge score strictly above 0.90.
Both checkpoints were trained for three epochs. This is not the RFT corpus.
Contents
data/train.jsonl retains nested messages, reasoning, tool calls and responses,
task prompts, rubric criteria and file specifications. Original… See the full description on the dataset page: https://huggingface.co/datasets/groundhogLLM/GraphForge-SFT-2169-clean.GraphRarebench
GraphRareBench Dataset
This release is the single frozen GraphRareBench benchmark described in the
AAAI 2027 submission. It contains 2,365 ontology-derived
rare-disease ranking cases and 18,093 labeled
target-confounder pairs.
Layout
dataset/
cases/graphrarebench_cases.jsonl
sidecars/evidence_bundle.jsonl
sidecars/graph_path_alignment.jsonl
metadata/manifest.json
metadata/release_summary.json
metadata/checksums.sha256
schema/case_public_schema.json… See the full description on the dataset page: https://huggingface.co/datasets/gcc009/GraphRarebench.discrete_prompting_scene_graphsapi_graph_reflectionprosqa_test_graph_4_coconut
