AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code
ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.
10168
1---2pretty_name: ArXiv Deep Learning Python Research Code3configs:4- config_name: default5 data_files:6 - split: train7 path: data/train-*8dataset_info:9 features:10 - name: repo11 dtype: string12 - name: file13 dtype: string14 - name: code15 dtype: string16 - name: file_length17 dtype: int6418 - name: avg_line_length19 dtype: float6420 - name: max_line_length21 dtype: int6422 - name: extension_type23 dtype: string24 splits:25 - name: train26 num_bytes: 3590067176.12519327 num_examples: 39149628 download_size: 149072432529 dataset_size: 3590067176.12519330language:31 - en32license: other33size_categories:34 - 100K<n<1M35tags:36 - code37 - deep-learning38 - arxiv39 - research40 - python41task_categories:42 - text-generation43---44 45# ArXiv Deep Learning Python Research Code46 47A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code.48 49## Dataset Summary50 51| Statistic | Value |52|-----------|-------|53| Total files | 391,496 |54| Total size | 1.49 GB |55| Source repos | 34,099 |56| Time span | ArXiv inception through July 2023 |57 58## Dataset Structure59 60| Field | Type | Description |61|-------|------|-------------|62| `repo` | string | GitHub repository name |63| `file` | string | File path in the repository |64| `code` | string | File contents |65| `file_length` | int64 | Number of characters in the file |66| `avg_line_length` | float64 | Average line length |67| `max_line_length` | int64 | Maximum line length |68| `extension_type` | string | File extension |69 70## Usage71 72```python73from datasets import load_dataset74 75# full dataset76ds = load_dataset("AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code", split="train")77 78# streaming79ds = load_dataset("AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code", streaming=True, split="train")80for sample in ds:81 print(sample["repo"], sample["file"])82 break83```84 85## Data Collection86 8734,099 active GitHub repository names were extracted from [ArXiv](https://arxiv.org/) papers from its inception through July 21st, 2023, totaling 773 GB of compressed GitHub repositories.88 89These repositories were filtered to files mentioning any of the following frameworks: `torch`, `jax`, `flax`, `stax`, `haiku`, `keras`, `fastai`, `xgboost`, `caffe`, `mxnet`, yielding 1.4 million files which were further filtered to the final 391k.90 91## Sensitive Information92 93The dataset may contain emails, IP addresses, and API/SSH keys that were previously published in public GitHub repositories.94 95## Related Resources96 97- [ArXiv DL Instruct](https://huggingface.co/datasets/AlgorithmicResearchGroup/ArXivDLInstruct) - Instruction-tuning dataset derived from this code98- [Algorithmic Research Group - Open Source](https://algorithmicresearchgroup.com/opensource.html)99 100## Citation101 102```bibtex103@misc{arxiv_deep_learning_python_research_code,104 title={ArXiv Deep Learning Python Research Code},105 author={Matthew Kenney},106 year={2023},107 publisher={Hugging Face},108 url={https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code}109}110```111 