Team Ai
Datasetpublic

cometadata/arxiv-software-repo-links

arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
0likes110downloads
Dataset Card

arXiv Software Repository Links

A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.

Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models

Quick Start

python
from datasets import load_dataset

# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")

# Load repo citation stats
repo_stats = load_dataset("cometadata/arxiv-software-repo-links", "repo_stats")

# Load user/org citation stats
user_stats = load_dataset("cometadata/arxiv-software-repo-links", "user_stats")

# Load co-citation data
repo_cocitations = load_dataset("cometadata/arxiv-software-repo-links", "repo_cocitations")
user_cocitations = load_dataset("cometadata/arxiv-software-repo-links", "user_cocitations")

# Load cluster data
repo_clusters = load_dataset("cometadata/arxiv-software-repo-links", "repo_clusters")
user_clusters = load_dataset("cometadata/arxiv-software-repo-links", "user_clusters")

Dataset Description

This dataset contains links between arXiv papers and software repositories (primarily GitHub), extracted from the full text and validated through multiple methods, including API checks and repository metadata analysis.

Configurations

ConfigDescriptionRecords
linksDOI-to-repository mappings with relationship type582,954
repo_statsAggregate citation counts per repository310,854
user_statsAggregate citation counts per GitHub user/organization163,474
repo_cocitationsRepository pairs cited together by the same papers745,602
user_cocitationsGitHub user/org pairs cited together by the same papers526,326
repo_clustersRepository communities detected via Louvain clustering219
user_clustersGitHub user/org communities detected via Louvain clustering38

Schema

links

json
{"doi": "10.48550/arxiv.2308.11197", "repo_url": "https://github.com/owner/repo", "relation_type": "References"}

repo_stats

json
{"repo_url": "https://github.com/owner/repo", "citation_count": 42}

user_stats

json
{"github_user": "facebookresearch", "citation_count": 3847}

repo_cocitations

json
{"repo_1": "https://github.com/google/jax", "repo_2": "https://github.com/google/flax", "cocitation_count": 144}

user_cocitations

json
{"github_user_1": "facebookresearch", "github_user_2": "microsoft", "cocitation_count": 84}

repo_clusters / user_clusters

json
{"cluster_id": 1, "size": 219, "top_members": ["kingoflolz/mesh-transformer-jax", "huggingface/trl", "..."], "members": ["..."]}

Relationship Types

TypeDescriptionCount
ReferencesPaper mentions/cites the repository447,807
IsSupplementedByRepository directly supports the paper (e.g., paper's code)135,147

Dataset Statistics

MetricValue
Total paper-repo links582,954
Unique repositories310,854
Unique GitHub users/orgs163,474
Unique arXiv DOIs343,537
Repo co-citation pairs745,602
User co-citation pairs526,326
Repo clusters219
User clusters38

Repository Citation Distribution

CitationsRepositories
1251,105
2-549,734
6-105,648
11-503,806
51-100333
100+228

Most Cited Repositories

Most Cited GitHub Users/Organizations

User/OrgCitations
facebookresearch8,599
huggingface6,148
google5,588
microsoft4,056
openai3,292
tatsu-lab3,192
google-research3,160
pytorch2,635
open-mmlab2,355
NVIDIA2,240

Top Co-citation Pairs (Repos)

Repo 1Repo 2Co-citations
google/flaxgoogle/jax200
MarekKowalski/FaceSwapdeepfakes/faceswap191
tatsu-lab/alpaca_evaltatsu-lab/stanford_alpaca179
ultralytics/ultralyticsultralytics/yolov5123
deepmind/dm-haikugoogle/jax114

Top Co-citation Pairs (Users/Orgs)

User 1User 2Co-citations
facebookresearchhuggingface322
huggingfacetatsu-lab300
facebookresearchmicrosoft284
facebookresearchpytorch281
facebookresearchgoogle-research245

Example Repository Clusters

ClusterThemeTop Members
1NLP/Transformershuggingface/transformers, pytorch/fairseq, google-research/bert
2LLMs/Instruction Tuningtatsu-lab/stanford_alpaca, kingoflolz/mesh-transformer-jax, huggingface/peft
3Computer Vision/Detectionfacebookresearch/detectron2, rwightman/pytorch-image-models, ultralytics/yolov5
4Generative Models/Diffusionblack-forest-labs/flux, huggingface/diffusers, openai/CLIP
5Deep Learning Frameworksfchollet/keras, tensorflow/models, pytorch/pytorch
6Reinforcement Learningopenai/baselines, hill-a/stable-baselines, DLR-RM/stable-baselines3
7Speech/Audiokaldi-asr/kaldi, espnet/espnet, openai/whisper

Example User/Org Clusters

ClusterThemeTop Members
1Astronomy/AstrophysicsLSSTDESC, dfm, jobovy, spacetelescope, cmbant
2Google/DeepMind Stackgoogle, deepmind, apache, facebook, jax-ml
3Computer Vision Researchfacebookresearch, open-mmlab, ultralytics, rwightman, Lightning-AI
4NLP/LLMshuggingface, allenai, UKPLab, hiyouga, castorini
5Google Research/Translationgoogle-research, fxsjy, moses-smt, mjpost, google-research-datasets
6Deep Learning Frameworkspytorch, tensorflow, fchollet, keras-team, fastai

Methodology

This dataset was produced using extract-link-arxiv-software-repos.

  1. 1.Software repository URLs extracted from the full text of arXiv works using pattern matching
  2. 2.URLs are then validated via git ls-remote, API checks, and HTTP verification
  3. 3.Generic "References" relationships are promoted to "IsSupplementedBy" when evidence suggests the repository directly supports the paper, based on:
  4. 4.arXiv ID present in repository README or description
  5. 5.Repository name similarity to paper title
  6. 6.Author matching, where the rrpository contributors matched to paper authors using the evamxb/dev-author-em-clf model from sci-soft-models
  7. 7.Co-citation constitutes pairs of repositories/users cited by the same paper are counted
  8. 8.For clustering, Louvain community detection us applied to co-citation networks to identify thematic clusters

License

CC0 1.0 Universal (Public Domain)

Citation

bibtex
@dataset{arxiv_software_repo_links,
  title={arXiv Software Repository Links},
  author={COMET},
  year={2025},
  publisher={Hugging Face},
  url={https://huggingface.co/datasets/cometadata/arxiv-software-repo-links}
}