datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
graph-pannuke
Graph-PanNuke: A Cell-Graph Dataset for Nucleus Classification from PanNuke
Graph-PanNuke is a node-level classification dataset derived from the PanNuke pan-cancer histology dataset. We use all slides at 40× magnification. Each tissue patch is converted into a cell-graph where nodes represent detected cell nuclei and edges encode spatial proximity. The task is predicting the cell type of each nucleus across 5 classes. Note that node features describe cell morphology, texture… See the full description on the dataset page: https://huggingface.co/datasets/ogutsevda/graph-pannuke.math-graph
Math-Graph
Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency
graph spanning both informal and formal mathematics. On the informal side it parses millions of
theorem-like environments from mathematics arXiv and recovers directed dependency edges within and
across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed
declaration dependencies across 25 Lean 4 projects. The two graphs are bridged into one… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/math-graph.Eedi-Misconceptions-Graph
Eedi Misconceptions Graph v1.0
A map of mathematical misconceptions and the curriculum constructs they appear in, released by Eedi under a CC BY 4.0 licence.
Eedi defines a misconception as a flawed conceptual structure, or a gap in conceptual understanding, that manifests as a systematic and predictable error pattern across problems involving the same mathematical concept. A construct is a small, specific element of mathematics — for example, "Order fractions with the same… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Eedi-Misconceptions-Graph.travel-fraud-graphs
TravelFraudBench (TFG)
The first publicly available labeled heterogeneous graph benchmark for GNN-based fraud ring detection in travel networks.
Paper
Dataset Structure
This dataset contains heterogeneous graph data split into 20 named configurations — one per node type and one per edge type — each with small, medium, and large splits.
Loading a specific node or edge table
from datasets import load_dataset
# Load user nodes (medium scale)
users =… See the full description on the dataset page: https://huggingface.co/datasets/bsajja7/travel-fraud-graphs.VG_scene_graph_cleanA cleaned version of Visual Genome dataset for textual scene graph parsing.
'has_attribute' relation is referred to as the relation 'is' in our processed data.
For details, please see "FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing", ACL 2023.
@article{li2023factual,
title={FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing},
author={Li, Zhuang and Chai, Yuyang and Zhuo, Terry Yue and Qu, Lizhen and Haffari, Gholamreza and Li, Fei… See the full description on the dataset page: https://huggingface.co/datasets/lizhuang144/VG_scene_graph_clean.math-graph
Math-Graph
Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency
graph spanning both informal and formal mathematics. On the informal side it parses millions of
theorem-like environments from mathematics arXiv and recovers directed dependency edges within and
across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed
declaration dependencies across 25 Lean 4 projects. The two graphs are bridged into one… See the full description on the dataset page: https://huggingface.co/datasets/jessteru/math-graph.FACTUAL_Scene_Graph_IDPlease refer to https://github.com/zhuang-li/FACTUAL for a detailed description of this dataset.
alanjo_graphics-card-full-specs
🟩NVIDIA & AMD🟥 GPUs Full Specs💠
Full specifications of NVIDIA and AMD graphics processing units (and more)
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 8,376
Files: 2
Files
gpu_specs_v6.csv
gpu_specs_v7.csv
Mirrored from Kaggle
sen_legal_graphrag_dataGraphCGGraphTranslator-arxivgraph-pannuke
Graph-PanNuke: A Cell-Graph Dataset for Nucleus Classification from PanNuke
Graph-PanNuke is a node-level classification dataset derived from the PanNuke pan-cancer histology dataset. We use all slides at 40× magnification. Each tissue patch is converted into a cell-graph where nodes represent detected cell nuclei and edges encode spatial proximity. The task is predicting the cell type of each nucleus across 5 classes. Note that node features describe cell morphology, texture… See the full description on the dataset page: https://huggingface.co/datasets/dszohib/graph-pannuke.dropthe-knowledge-graph
DropThe Entity Relationship Graph
A large-scale entity relationship dataset containing 2.9 million typed, directional connections between 1.8 million entities spanning entertainment, media, finance, and technology. Extracted from the DropThe knowledge graph.
Dataset Description
While most open knowledge graphs focus on encyclopedic facts (Wikidata) or narrow domains (MovieLens for ratings), this dataset captures operational relationships -- the connections that actually… See the full description on the dataset page: https://huggingface.co/datasets/DropTheHQ/dropthe-knowledge-graph.g-retriever-scene-graphsknowledge-graph
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset can be used for training the SVO
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/Aditya210/knowledge-graph.clinical-nearmiss-counterfactual-outcome-divergence-graphing-v0.1What this dataset tests
Whether a model can describe where clinical trajectories divergebetween a failing path and a repaired path.
Required outputs
divergence_point
avoided_failure_mode
minimal_intervention_required
Typical failures
describing only the repaired path
missing the divergence point
proposing a large intervention when a small one would suffice
Suggested prompt wrapper
System
You graph counterfactual outcome divergence.
User
Failing path… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-nearmiss-counterfactual-outcome-divergence-graphing-v0.1.2d-graphene-conductivity-coherence-loss-v0.1Goal
Predict irreversible conductivity dropin graphene devices.
Core idea
Failure does not arrive as one threshold.
It arrives when critical relationships collapse:
Raman D/G ratiomobilitythermal conductivitysheet resistance
stop telling one coherent story.
Inputs
operational hours
Raman D/G ratio
carrier mobility
thermal conductivity
sheet resistance
defect density
humidity exposure
Required outputs
conductivity_coherence_score
decoupling_flag
decoupling_type
irreversible_drop_probability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/2d-graphene-conductivity-coherence-loss-v0.1.legal-citation-graph
DEPRECATED: Use the new dataset instead https://huggingface.co/datasets/openlegaldata/legal-citation-graph-germany
Citation graph for German court decisions based on https://openlegaldata.io/research/2019/02/19/court-decision-dataset.html
More recent data and a better HF integration soon!
Graph2Text_rdf_typegraph-of-thoughts-sortingSource Paper: https://arxiv.org/abs/2308.09687
Citation
@article{Besta_2024,
title={Graph of Thoughts: Solving Elaborate Problems with Large Language Models},
volume={38},
ISSN={2159-5399},
url={http://dx.doi.org/10.1609/aaai.v38i16.29720},
DOI={10.1609/aaai.v38i16.29720},
number={16},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
publisher={Association for the Advancement of Artificial Intelligence (AAAI)},
author={Besta, Maciej… See the full description on the dataset page: https://huggingface.co/datasets/SauravMaheshkar/graph-of-thoughts-sorting.Grapheme128x128
This data set is preprocess version of this competition data set. The preprocess data is collected from here.
Japanese-Complex-GraphQA
Japanese-Complex-GraphQA
Japanese-Complex-GraphQA は,日本語の図表(チャート・グラフ)を対象とした図表質問応答(Graph Question Answering)データセットです。高度な読解・比較・計算を必要とする難易度の高い質問を含むことを特徴としています。
本データセットは,図表を含む公開統計資料を基に,質問文おび正答を新たに作成し,日本語環境におけるマルチモーダル QA および推論能力の評価を目的として構築されました。
データセット構成
本データセットは以下のファイルから構成されます。
ファイル一覧
jc_graphqa.csv質問応答データ(id, 質問文, 正答, 画像ファイルid)
jc_graphqa_tags.jsonl各質問に付与された推論・図表属性タグ
CSV ファイル形式(data.csv)
カラム名
内容
id
データ識別子(数値)
question
日本語の質問文
answer
正答… See the full description on the dataset page: https://huggingface.co/datasets/ab528/Japanese-Complex-GraphQA.devign_for_graphormerFACTUAL_Scene_GraphThe scene graph parsing dataset described in FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing
Please see details from https://github.com/zhuang-li/FACTUAL .
graph_shortest_path_n15_d2-8graph_shortest_path_n10_d2-8graph_shortest_path_n30_d6-10cricketstudio-knowledge-graph
CricketStudio Knowledge Graph — Sample
A sample of the CricketStudio cricket knowledge graph: the Royal Challengers
Bengaluru squad and their batter-vs-bowler matchup relationships, from IPL 2026
plus IPL-career head-to-head data. Every entity links back to its canonical page
on https://players.cricketstudio.ai via canonical_url.
Files / configs
File
Rows
What
nodes.csv / nodes.json
141
entities — id (slug), type, name, canonical_url
edges.csv /… See the full description on the dataset page: https://huggingface.co/datasets/CricketStudio/cricketstudio-knowledge-graph.graph_shortest_path_n20_d6-10graph_shortest_path_n12_d2-8
