Team Ai
Datasetpublic

AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code

ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.

sourceHugging Faceotherupdated 6mo agoView on Hugging Face
10likes179downloads
Dataset Card

ArXiv Deep Learning Python Research Code

A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code.

Dataset Summary

StatisticValue
Total files391,496
Total size1.49 GB
Source repos34,099
Time spanArXiv inception through July 2023

Dataset Structure

FieldTypeDescription
repostringGitHub repository name
filestringFile path in the repository
codestringFile contents
file_lengthint64Number of characters in the file
avg_line_lengthfloat64Average line length
max_line_lengthint64Maximum line length
extension_typestringFile extension

Usage

python
from datasets import load_dataset

# full dataset
ds = load_dataset("AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code", split="train")

# streaming
ds = load_dataset("AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code", streaming=True, split="train")
for sample in ds:
    print(sample["repo"], sample["file"])
    break

Data Collection

34,099 active GitHub repository names were extracted from ArXiv papers from its inception through July 21st, 2023, totaling 773 GB of compressed GitHub repositories.

These repositories were filtered to files mentioning any of the following frameworks: torch, jax, flax, stax, haiku, keras, fastai, xgboost, caffe, mxnet, yielding 1.4 million files which were further filtered to the final 391k.

Sensitive Information

The dataset may contain emails, IP addresses, and API/SSH keys that were previously published in public GitHub repositories.

Related Resources

Citation

bibtex
@misc{arxiv_deep_learning_python_research_code,
    title={ArXiv Deep Learning Python Research Code},
    author={Matthew Kenney},
    year={2023},
    publisher={Hugging Face},
    url={https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code}
}