cometadata/arxiv-software-repo-links
arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
# Load repo citation stats
repo_stats = load_dataset("cometadata/arxiv-software-repo-links", "repo_stats")
# Load user/org citation stats
user_stats = load_dataset("cometadata/arxiv-software-repo-links", "user_stats")
# Load co-citation data
repo_cocitations = load_dataset("cometadata/arxiv-software-repo-links", "repo_cocitations")
user_cocitations = load_dataset("cometadata/arxiv-software-repo-links", "user_cocitations")
# Load cluster data
repo_clusters = load_dataset("cometadata/arxiv-software-repo-links", "repo_clusters")
user_clusters = load_dataset("cometadata/arxiv-software-repo-links", "user_clusters")Dataset Description
This dataset contains links between arXiv papers and software repositories (primarily GitHub), extracted from the full text and validated through multiple methods, including API checks and repository metadata analysis.
Configurations
Schema
links
{"doi": "10.48550/arxiv.2308.11197", "repo_url": "https://github.com/owner/repo", "relation_type": "References"}repo_stats
{"repo_url": "https://github.com/owner/repo", "citation_count": 42}user_stats
{"github_user": "facebookresearch", "citation_count": 3847}repo_cocitations
{"repo_1": "https://github.com/google/jax", "repo_2": "https://github.com/google/flax", "cocitation_count": 144}user_cocitations
{"github_user_1": "facebookresearch", "github_user_2": "microsoft", "cocitation_count": 84}repo_clusters / user_clusters
{"cluster_id": 1, "size": 219, "top_members": ["kingoflolz/mesh-transformer-jax", "huggingface/trl", "..."], "members": ["..."]}Relationship Types
Dataset Statistics
Repository Citation Distribution
Most Cited Repositories
Most Cited GitHub Users/Organizations
Top Co-citation Pairs (Repos)
Top Co-citation Pairs (Users/Orgs)
Example Repository Clusters
Example User/Org Clusters
Methodology
This dataset was produced using extract-link-arxiv-software-repos.
- Software repository URLs extracted from the full text of arXiv works using pattern matching
- URLs are then validated via git ls-remote, API checks, and HTTP verification
- Generic "References" relationships are promoted to "IsSupplementedBy" when evidence suggests the repository directly supports the paper, based on:
- arXiv ID present in repository README or description
- Repository name similarity to paper title
- Author matching, where the rrpository contributors matched to paper authors using the evamxb/dev-author-em-clf model from sci-soft-models
- Co-citation constitutes pairs of repositories/users cited by the same paper are counted
- For clustering, Louvain community detection us applied to co-citation networks to identify thematic clusters
License
CC0 1.0 Universal (Public Domain)
Citation
@dataset{arxiv_software_repo_links,
title={arXiv Software Repository Links},
author={COMET},
year={2025},
publisher={Hugging Face},
url={https://huggingface.co/datasets/cometadata/arxiv-software-repo-links}
}