Team Ai
Datasetpublic

AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code

ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.

sourceHugging Faceotherupdated 6mo agoView on Hugging Face
10likes168downloads
README.md111 linesDownload Raw Back to root
1---2pretty_name: ArXiv Deep Learning Python Research Code3configs:4- config_name: default5  data_files:6  - split: train7    path: data/train-*8dataset_info:9  features:10  - name: repo11    dtype: string12  - name: file13    dtype: string14  - name: code15    dtype: string16  - name: file_length17    dtype: int6418  - name: avg_line_length19    dtype: float6420  - name: max_line_length21    dtype: int6422  - name: extension_type23    dtype: string24  splits:25  - name: train26    num_bytes: 3590067176.12519327    num_examples: 39149628  download_size: 149072432529  dataset_size: 3590067176.12519330language:31  - en32license: other33size_categories:34  - 100K<n<1M35tags:36  - code37  - deep-learning38  - arxiv39  - research40  - python41task_categories:42  - text-generation43---44 45# ArXiv Deep Learning Python Research Code46 47A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code.48 49## Dataset Summary50 51| Statistic | Value |52|-----------|-------|53| Total files | 391,496 |54| Total size | 1.49 GB |55| Source repos | 34,099 |56| Time span | ArXiv inception through July 2023 |57 58## Dataset Structure59 60| Field | Type | Description |61|-------|------|-------------|62| `repo` | string | GitHub repository name |63| `file` | string | File path in the repository |64| `code` | string | File contents |65| `file_length` | int64 | Number of characters in the file |66| `avg_line_length` | float64 | Average line length |67| `max_line_length` | int64 | Maximum line length |68| `extension_type` | string | File extension |69 70## Usage71 72```python73from datasets import load_dataset74 75# full dataset76ds = load_dataset("AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code", split="train")77 78# streaming79ds = load_dataset("AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code", streaming=True, split="train")80for sample in ds:81    print(sample["repo"], sample["file"])82    break83```84 85## Data Collection86 8734,099 active GitHub repository names were extracted from [ArXiv](https://arxiv.org/) papers from its inception through July 21st, 2023, totaling 773 GB of compressed GitHub repositories.88 89These repositories were filtered to files mentioning any of the following frameworks: `torch`, `jax`, `flax`, `stax`, `haiku`, `keras`, `fastai`, `xgboost`, `caffe`, `mxnet`, yielding 1.4 million files which were further filtered to the final 391k.90 91## Sensitive Information92 93The dataset may contain emails, IP addresses, and API/SSH keys that were previously published in public GitHub repositories.94 95## Related Resources96 97- [ArXiv DL Instruct](https://huggingface.co/datasets/AlgorithmicResearchGroup/ArXivDLInstruct) - Instruction-tuning dataset derived from this code98- [Algorithmic Research Group - Open Source](https://algorithmicresearchgroup.com/opensource.html)99 100## Citation101 102```bibtex103@misc{arxiv_deep_learning_python_research_code,104    title={ArXiv Deep Learning Python Research Code},105    author={Matthew Kenney},106    year={2023},107    publisher={Hugging Face},108    url={https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code}109}110```111