Team Ai
Datasetpublic

jugalgajjar/CS-Knowledge-Graph-Dataset

CS Knowledge Graph Dataset A multi-scale heterogeneous knowledge graph of Computer Science scholarly data, built from OpenAlex. Each scale is an independent, self-contained subgraph centered on Computer Science papers, their authors, publication venues, and concept tags, plus the relationships between them. The dataset is intended for research on knowledge graph embeddings, link prediction, node classification, scholarly recommendation, and graph neural networks at varying… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/CS-Knowledge-Graph-Dataset.

sourceHugging Facecc-by-sa-4.0updated 5mo agoView on Hugging Face
0likes97downloads
README.md286 linesDownload Raw Back to root
1---2license: cc-by-sa-4.03language:4- en5pretty_name: CS Knowledge Graph (OpenAlex)6size_categories:7- 10M<n<100M8task_categories:9- graph-ml10- feature-extraction11tags:12- knowledge-graph13- openalex14- computer-science15- bibliographic16- citation-network17- co-authorship18- scholarly19- link-prediction20- node-classification21configs:22- config_name: 1k_nodes23  default: true24  data_files:25  - split: train26    path: 1k/nodes.parquet27- config_name: 1k_edges28  data_files:29  - split: train30    path: 1k/edges.parquet31- config_name: 10k_nodes32  data_files:33  - split: train34    path: 10k/nodes.parquet35- config_name: 10k_edges36  data_files:37  - split: train38    path: 10k/edges.parquet39- config_name: 100k_nodes40  data_files:41  - split: train42    path: 100k/nodes.parquet43- config_name: 100k_edges44  data_files:45  - split: train46    path: 100k/edges.parquet47- config_name: 1m_nodes48  data_files:49  - split: train50    path: 1m/nodes.parquet51- config_name: 1m_edges52  data_files:53  - split: train54    path: 1m/edges.parquet55- config_name: 10m_nodes56  data_files:57  - split: train58    path: 10m/nodes.parquet59- config_name: 10m_edges60  data_files:61  - split: train62    path: 10m/edges.parquet63---64 65# CS Knowledge Graph Dataset66 67[![License](https://img.shields.io/badge/License-CC%20BY--SA%204.0-lightgrey.svg)](https://creativecommons.org/licenses/by-sa/4.0/)68[![GitHub](https://img.shields.io/badge/GitHub-Repository-181717.svg?logo=github)](https://github.com/JugalGajjar/HyperComplEx-Multi-Space-KG-Embeddings)69[![IEEE Xplore](https://img.shields.io/badge/IEEE_Xplore-11400828-00629B.svg)](https://ieeexplore.ieee.org/abstract/document/11400828)70 71A multi-scale heterogeneous knowledge graph of Computer Science scholarly data,72built from [OpenAlex](https://openalex.org). Each scale is an independent,73self-contained subgraph centered on Computer Science papers, their authors,74publication venues, and concept tags, plus the relationships between them.75 76The dataset is intended for research on knowledge graph embeddings, link77prediction, node classification, scholarly recommendation, and graph neural78networks at varying scales of compute.79 80## Scales81 82Five scales are provided so the same pipeline can be benchmarked from quick83prototyping (1k) to large-scale training (10m). Each scale is a strict superset84of the smaller ones in spirit, but is sampled independently — treat them as85five separate graphs rather than nested cuts.86 87| Config | Nodes      | Edges       | Parquet size | Raw SQLite (zip) |88|--------|-----------:|------------:|-------------:|-----------------:|89| `1k`   |      5,237 |      32,655 |       277 KB |           961 KB |90| `10k`  |     44,933 |     252,631 |       2.0 MB |           7.7 MB |91| `100k` |    348,983 |   2,162,386 |        16 MB |            68 MB |92| `1m`   |  2,384,896 |  13,530,177 |       117 MB |           597 MB |93| `10m`  |  7,210,506 |  44,631,484 |       384 MB |           2.1 GB |94 95## Schema96 97Each scale exposes two configs, `<scale>_nodes` and `<scale>_edges`. They98share a single split named `train` (a `datasets` convention — there is no99held-out test split, since the intended use is to define your own splits over100the graph).101 102### `nodes` config103 104| Column       | Type   | Description                                                           |105|--------------|--------|-----------------------------------------------------------------------|106| `node_id`    | string | Unique node identifier, prefixed by type (e.g. `paper_W2604738573`).  |107| `node_name`  | string | Human-readable name (paper title, author display name, venue, etc.).  |108| `node_type`  | string | One of `Paper`, `Author`, `Venue`, `Concept`.                         |109| `attributes` | string | Type-specific attributes encoded as a JSON string (see below).        |110 111The `attributes` JSON object has different keys depending on `node_type`:112 113- **Paper**: `year` (int), `citation_count` (int), `venue` (string), `type` (string, e.g. `article`)114- **Author**: `h_index` (int or null), `citation_count` (int or null), `works_count` (int or null), `institution` (string)115- **Venue**: `type` (string, e.g. `journal`, `conference`), `publisher` (string)116- **Concept**: `domain` (string, e.g. `CS`)117 118### `edges` config119 120| Column     | Type   | Description                                                                                |121|------------|--------|--------------------------------------------------------------------------------------------|122| `source`   | string | `node_id` of the source node.                                                              |123| `relation` | string | One of `AUTHORED`, `CITES`, `PUBLISHED_IN`, `BELONGS_TO`, `COLLABORATES_WITH`.             |124| `target`   | string | `node_id` of the target node.                                                              |125| `year`     | float  | Year associated with the edge when applicable (e.g. publication year); `null` otherwise.   |126 127Relation semantics:128 129- `AUTHORED` — `Author → Paper`130- `CITES` — `Paper → Paper`131- `PUBLISHED_IN` — `Paper → Venue`132- `BELONGS_TO` — `Paper → Concept`133- `COLLABORATES_WITH` — `Author → Author` (co-authorship; symmetric, may appear in both directions)134 135**Dangling `CITES` targets.** Each scale is built from a Computer Science slice136of OpenAlex, so the `nodes` table only contains CS papers (plus their authors,137venues, and concepts). However, those CS papers may cite papers from outside138CS — those external papers appear as `target` in `CITES` edges but are **not**139present in the `nodes` table. Filter or add placeholder nodes as appropriate140for your task. Sources are always present in `nodes`; only `CITES` targets can141be dangling.142 143## Usage144 145### Load with the `datasets` library146 147```python148from datasets import load_dataset149 150# Configs follow the pattern "<scale>_nodes" / "<scale>_edges".151# Scales: 1k, 10k, 100k, 1m, 10m152nodes = load_dataset("jugalgajjar/CS-Knowledge-Graph-Dataset", "10k_nodes", split="train")153edges = load_dataset("jugalgajjar/CS-Knowledge-Graph-Dataset", "10k_edges", split="train")154 155print(nodes[0])156# {'node_id': 'paper_W...', 'node_name': '...', 'node_type': 'Paper',157#  'attributes': '{"year": 2016, "citation_count": 1816, ...}'}158 159import json160attrs = json.loads(nodes[0]["attributes"])161```162 163### Load directly with pandas / pyarrow164 165```python166import pandas as pd167nodes = pd.read_parquet("hf://datasets/jugalgajjar/CS-Knowledge-Graph-Dataset/100k/nodes.parquet")168edges = pd.read_parquet("hf://datasets/jugalgajjar/CS-Knowledge-Graph-Dataset/100k/edges.parquet")169```170 171### Build a PyTorch Geometric graph172 173```python174import numpy as np175import torch176from torch_geometric.data import HeteroData177from datasets import load_dataset178 179scale = "10k"180nodes = load_dataset("jugalgajjar/CS-Knowledge-Graph-Dataset", f"{scale}_nodes", split="train").to_pandas()181edges = load_dataset("jugalgajjar/CS-Knowledge-Graph-Dataset", f"{scale}_edges", split="train").to_pandas()182 183# Build per-type id -> contiguous index maps184data = HeteroData()185id_maps = {}186for ntype, group in nodes.groupby("node_type"):187    ids = group["node_id"].tolist()188    id_maps[ntype] = {nid: i for i, nid in enumerate(ids)}189    data[ntype].num_nodes = len(ids)190 191# Each node_id is prefixed with its type192type_from_prefix = {"paper": "Paper", "author": "Author", "venue": "Venue", "concept": "Concept"}193def ntype_of(nid: str) -> str:194    return type_from_prefix[nid.split("_", 1)[0]]195 196# Drop CITES edges whose target isn't in the node set (cross-domain citations).197node_id_set = set(nodes["node_id"])198edges = edges[edges["target"].isin(node_id_set)].reset_index(drop=True)199 200for relation, group in edges.groupby("relation"):201    src_type = ntype_of(group["source"].iloc[0])202    dst_type = ntype_of(group["target"].iloc[0])203    src = group["source"].map(id_maps[src_type]).to_numpy(dtype=np.int64)204    dst = group["target"].map(id_maps[dst_type]).to_numpy(dtype=np.int64)205    data[src_type, relation, dst_type].edge_index = torch.from_numpy(np.stack([src, dst]))206 207print(data)208```209 210## Raw SQLite databases211 212In addition to the Parquet files, the original SQLite databases used to build213each scale are available under `raw/`:214 215```216raw/cs1k_openalex.db.zip217raw/cs10k_openalex.db.zip218raw/cs100k_openalex.db.zip219raw/cs1m_openalex.db.zip220raw/cs10m_openalex.db.zip221```222 223These are useful if you want to run SQL queries over the source records224directly. Download with `huggingface_hub`:225 226```python227from huggingface_hub import hf_hub_download228path = hf_hub_download(229    repo_id="jugalgajjar/CS-Knowledge-Graph-Dataset",230    repo_type="dataset",231    filename="raw/cs10k_openalex.db.zip",232)233```234 235## Citation236 237This dataset was introduced in the following paper. **If you use this dataset238in your work, please cite it.** Please also cite OpenAlex (the source data;239see their [citation guidance](https://docs.openalex.org)).240 241**BibTeX:**242 243```bibtex244@inproceedings{gajjar2025hypercomplex,245  title={HyperComplEx: Adaptive Multi-Space Knowledge Graph Embeddings},246  author={Gajjar, Jugal and Ranaware, Kaustik and Subramaniakuppusamy, Kamalasankari and Gandhi, Vaibhav C},247  booktitle={2025 IEEE International Conference on Big Data (BigData)},248  pages={5623--5631},249  year={2025},250  organization={IEEE}251}252```253 254**APA:**255 256> Gajjar, J., Ranaware, K., Subramaniakuppusamy, K., & Gandhi, V. C. (2025, December). HyperComplEx: Adaptive Multi-Space Knowledge Graph Embeddings. In *2025 IEEE International Conference on Big Data (BigData)* (pp. 5623–5631). IEEE.257 258## Source and licensing259 260- **Source data:** [OpenAlex](https://openalex.org), released into the public261  domain under [CC0](https://creativecommons.org/publicdomain/zero/1.0/).262- **This derived dataset:** licensed under263  [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). You may use,264  modify, and redistribute it, including commercially, provided you give265  attribution and license your derivative works under the same terms.266 267## Repository layout268 269```270.271├── README.md272├── 1k/273│   ├── nodes.parquet274│   └── edges.parquet275├── 10k/    (same layout)276├── 100k/   (same layout)277├── 1m/     (same layout)278├── 10m/    (same layout)279└── raw/280    ├── cs1k_openalex.db.zip281    ├── cs10k_openalex.db.zip282    ├── cs100k_openalex.db.zip283    ├── cs1m_openalex.db.zip284    └── cs10m_openalex.db.zip285```286