jugalgajjar/CS-Knowledge-Graph-Dataset
CS Knowledge Graph Dataset A multi-scale heterogeneous knowledge graph of Computer Science scholarly data, built from OpenAlex. Each scale is an independent, self-contained subgraph centered on Computer Science papers, their authors, publication venues, and concept tags, plus the relationships between them. The dataset is intended for research on knowledge graph embeddings, link prediction, node classification, scholarly recommendation, and graph neural networks at varying… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/CS-Knowledge-Graph-Dataset.
097
1---2license: cc-by-sa-4.03language:4- en5pretty_name: CS Knowledge Graph (OpenAlex)6size_categories:7- 10M<n<100M8task_categories:9- graph-ml10- feature-extraction11tags:12- knowledge-graph13- openalex14- computer-science15- bibliographic16- citation-network17- co-authorship18- scholarly19- link-prediction20- node-classification21configs:22- config_name: 1k_nodes23 default: true24 data_files:25 - split: train26 path: 1k/nodes.parquet27- config_name: 1k_edges28 data_files:29 - split: train30 path: 1k/edges.parquet31- config_name: 10k_nodes32 data_files:33 - split: train34 path: 10k/nodes.parquet35- config_name: 10k_edges36 data_files:37 - split: train38 path: 10k/edges.parquet39- config_name: 100k_nodes40 data_files:41 - split: train42 path: 100k/nodes.parquet43- config_name: 100k_edges44 data_files:45 - split: train46 path: 100k/edges.parquet47- config_name: 1m_nodes48 data_files:49 - split: train50 path: 1m/nodes.parquet51- config_name: 1m_edges52 data_files:53 - split: train54 path: 1m/edges.parquet55- config_name: 10m_nodes56 data_files:57 - split: train58 path: 10m/nodes.parquet59- config_name: 10m_edges60 data_files:61 - split: train62 path: 10m/edges.parquet63---64 65# CS Knowledge Graph Dataset66 67[](https://creativecommons.org/licenses/by-sa/4.0/)68[](https://github.com/JugalGajjar/HyperComplEx-Multi-Space-KG-Embeddings)69[](https://ieeexplore.ieee.org/abstract/document/11400828)70 71A multi-scale heterogeneous knowledge graph of Computer Science scholarly data,72built from [OpenAlex](https://openalex.org). Each scale is an independent,73self-contained subgraph centered on Computer Science papers, their authors,74publication venues, and concept tags, plus the relationships between them.75 76The dataset is intended for research on knowledge graph embeddings, link77prediction, node classification, scholarly recommendation, and graph neural78networks at varying scales of compute.79 80## Scales81 82Five scales are provided so the same pipeline can be benchmarked from quick83prototyping (1k) to large-scale training (10m). Each scale is a strict superset84of the smaller ones in spirit, but is sampled independently — treat them as85five separate graphs rather than nested cuts.86 87| Config | Nodes | Edges | Parquet size | Raw SQLite (zip) |88|--------|-----------:|------------:|-------------:|-----------------:|89| `1k` | 5,237 | 32,655 | 277 KB | 961 KB |90| `10k` | 44,933 | 252,631 | 2.0 MB | 7.7 MB |91| `100k` | 348,983 | 2,162,386 | 16 MB | 68 MB |92| `1m` | 2,384,896 | 13,530,177 | 117 MB | 597 MB |93| `10m` | 7,210,506 | 44,631,484 | 384 MB | 2.1 GB |94 95## Schema96 97Each scale exposes two configs, `<scale>_nodes` and `<scale>_edges`. They98share a single split named `train` (a `datasets` convention — there is no99held-out test split, since the intended use is to define your own splits over100the graph).101 102### `nodes` config103 104| Column | Type | Description |105|--------------|--------|-----------------------------------------------------------------------|106| `node_id` | string | Unique node identifier, prefixed by type (e.g. `paper_W2604738573`). |107| `node_name` | string | Human-readable name (paper title, author display name, venue, etc.). |108| `node_type` | string | One of `Paper`, `Author`, `Venue`, `Concept`. |109| `attributes` | string | Type-specific attributes encoded as a JSON string (see below). |110 111The `attributes` JSON object has different keys depending on `node_type`:112 113- **Paper**: `year` (int), `citation_count` (int), `venue` (string), `type` (string, e.g. `article`)114- **Author**: `h_index` (int or null), `citation_count` (int or null), `works_count` (int or null), `institution` (string)115- **Venue**: `type` (string, e.g. `journal`, `conference`), `publisher` (string)116- **Concept**: `domain` (string, e.g. `CS`)117 118### `edges` config119 120| Column | Type | Description |121|------------|--------|--------------------------------------------------------------------------------------------|122| `source` | string | `node_id` of the source node. |123| `relation` | string | One of `AUTHORED`, `CITES`, `PUBLISHED_IN`, `BELONGS_TO`, `COLLABORATES_WITH`. |124| `target` | string | `node_id` of the target node. |125| `year` | float | Year associated with the edge when applicable (e.g. publication year); `null` otherwise. |126 127Relation semantics:128 129- `AUTHORED` — `Author → Paper`130- `CITES` — `Paper → Paper`131- `PUBLISHED_IN` — `Paper → Venue`132- `BELONGS_TO` — `Paper → Concept`133- `COLLABORATES_WITH` — `Author → Author` (co-authorship; symmetric, may appear in both directions)134 135**Dangling `CITES` targets.** Each scale is built from a Computer Science slice136of OpenAlex, so the `nodes` table only contains CS papers (plus their authors,137venues, and concepts). However, those CS papers may cite papers from outside138CS — those external papers appear as `target` in `CITES` edges but are **not**139present in the `nodes` table. Filter or add placeholder nodes as appropriate140for your task. Sources are always present in `nodes`; only `CITES` targets can141be dangling.142 143## Usage144 145### Load with the `datasets` library146 147```python148from datasets import load_dataset149 150# Configs follow the pattern "<scale>_nodes" / "<scale>_edges".151# Scales: 1k, 10k, 100k, 1m, 10m152nodes = load_dataset("jugalgajjar/CS-Knowledge-Graph-Dataset", "10k_nodes", split="train")153edges = load_dataset("jugalgajjar/CS-Knowledge-Graph-Dataset", "10k_edges", split="train")154 155print(nodes[0])156# {'node_id': 'paper_W...', 'node_name': '...', 'node_type': 'Paper',157# 'attributes': '{"year": 2016, "citation_count": 1816, ...}'}158 159import json160attrs = json.loads(nodes[0]["attributes"])161```162 163### Load directly with pandas / pyarrow164 165```python166import pandas as pd167nodes = pd.read_parquet("hf://datasets/jugalgajjar/CS-Knowledge-Graph-Dataset/100k/nodes.parquet")168edges = pd.read_parquet("hf://datasets/jugalgajjar/CS-Knowledge-Graph-Dataset/100k/edges.parquet")169```170 171### Build a PyTorch Geometric graph172 173```python174import numpy as np175import torch176from torch_geometric.data import HeteroData177from datasets import load_dataset178 179scale = "10k"180nodes = load_dataset("jugalgajjar/CS-Knowledge-Graph-Dataset", f"{scale}_nodes", split="train").to_pandas()181edges = load_dataset("jugalgajjar/CS-Knowledge-Graph-Dataset", f"{scale}_edges", split="train").to_pandas()182 183# Build per-type id -> contiguous index maps184data = HeteroData()185id_maps = {}186for ntype, group in nodes.groupby("node_type"):187 ids = group["node_id"].tolist()188 id_maps[ntype] = {nid: i for i, nid in enumerate(ids)}189 data[ntype].num_nodes = len(ids)190 191# Each node_id is prefixed with its type192type_from_prefix = {"paper": "Paper", "author": "Author", "venue": "Venue", "concept": "Concept"}193def ntype_of(nid: str) -> str:194 return type_from_prefix[nid.split("_", 1)[0]]195 196# Drop CITES edges whose target isn't in the node set (cross-domain citations).197node_id_set = set(nodes["node_id"])198edges = edges[edges["target"].isin(node_id_set)].reset_index(drop=True)199 200for relation, group in edges.groupby("relation"):201 src_type = ntype_of(group["source"].iloc[0])202 dst_type = ntype_of(group["target"].iloc[0])203 src = group["source"].map(id_maps[src_type]).to_numpy(dtype=np.int64)204 dst = group["target"].map(id_maps[dst_type]).to_numpy(dtype=np.int64)205 data[src_type, relation, dst_type].edge_index = torch.from_numpy(np.stack([src, dst]))206 207print(data)208```209 210## Raw SQLite databases211 212In addition to the Parquet files, the original SQLite databases used to build213each scale are available under `raw/`:214 215```216raw/cs1k_openalex.db.zip217raw/cs10k_openalex.db.zip218raw/cs100k_openalex.db.zip219raw/cs1m_openalex.db.zip220raw/cs10m_openalex.db.zip221```222 223These are useful if you want to run SQL queries over the source records224directly. Download with `huggingface_hub`:225 226```python227from huggingface_hub import hf_hub_download228path = hf_hub_download(229 repo_id="jugalgajjar/CS-Knowledge-Graph-Dataset",230 repo_type="dataset",231 filename="raw/cs10k_openalex.db.zip",232)233```234 235## Citation236 237This dataset was introduced in the following paper. **If you use this dataset238in your work, please cite it.** Please also cite OpenAlex (the source data;239see their [citation guidance](https://docs.openalex.org)).240 241**BibTeX:**242 243```bibtex244@inproceedings{gajjar2025hypercomplex,245 title={HyperComplEx: Adaptive Multi-Space Knowledge Graph Embeddings},246 author={Gajjar, Jugal and Ranaware, Kaustik and Subramaniakuppusamy, Kamalasankari and Gandhi, Vaibhav C},247 booktitle={2025 IEEE International Conference on Big Data (BigData)},248 pages={5623--5631},249 year={2025},250 organization={IEEE}251}252```253 254**APA:**255 256> Gajjar, J., Ranaware, K., Subramaniakuppusamy, K., & Gandhi, V. C. (2025, December). HyperComplEx: Adaptive Multi-Space Knowledge Graph Embeddings. In *2025 IEEE International Conference on Big Data (BigData)* (pp. 5623–5631). IEEE.257 258## Source and licensing259 260- **Source data:** [OpenAlex](https://openalex.org), released into the public261 domain under [CC0](https://creativecommons.org/publicdomain/zero/1.0/).262- **This derived dataset:** licensed under263 [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). You may use,264 modify, and redistribute it, including commercially, provided you give265 attribution and license your derivative works under the same terms.266 267## Repository layout268 269```270.271├── README.md272├── 1k/273│ ├── nodes.parquet274│ └── edges.parquet275├── 10k/ (same layout)276├── 100k/ (same layout)277├── 1m/ (same layout)278├── 10m/ (same layout)279└── raw/280 ├── cs1k_openalex.db.zip281 ├── cs10k_openalex.db.zip282 ├── cs100k_openalex.db.zip283 ├── cs1m_openalex.db.zip284 └── cs10m_openalex.db.zip285```286 