Team Ai
Datasetpublic

CodedotAI/code_clippy_github

The Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.

sourceHugging Facemitupdated 4y agoView on Hugging Face
20likes3.4kdownloads
README.md176 linesDownload Raw Back to root
1---2annotations_creators: []3language_creators:4- crowdsourced5- expert-generated6language: ["code"]7license:8- mit9multilinguality:10- multilingual11pretty_name: code-clippy-github-code12size_categories:13- unknown14source_datasets: []15task_categories:16- sequence-modeling17task_ids:18- language-modeling19---20# Code Clippy Github Dataset21## Dataset Description22The Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totaling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BigQuery.23### How to use it24This dataset is pretty large please use the streaming parameter from the  `datasets` library as seen below:25```python26from datasets import load_dataset27 28ds = load_dataset("CodedotAI/code_clippy_github", streaming=True)29 30```31## Data Structure32 33### Data Instances34 35```python36{37 'code_text': " a = mc^2",38 'repo_name': 'NotEinstein',39 'file_path': 'root/users/einstein.py',40 'language': 'Python',41 'license': 'isc',42 'size': 243} 44```45 46### Data Fields47 48|Field|Type|Description|49|---|---|---|50|code_text|string|string of the source code contained in the code file|51|repo_name|string|name of the GitHub repository|52|file_path|string|path of the code file within the repository |53|language|string|programming language used in the file inferred by the file extension|54|license|string|license of GitHub repository|55|size|int|size of source file in bytes|56### Data Splits57Only a train split is provided in this dataset.58## Languages59The dataset contains 22 programming languages with over 23 extensions:60```python61{62    "C": [".c"],63    "C#": [".cs"],64    "C++": [".cpp"],65    "CSS": [".css"],66    "Dart" : [".dart"],67    "GO": [".go"],68    "HTML":[".html"],69    "Java": [".java"],70    "JavaScript": [".js"],71    "Jupyter Notebooks (Python)": [".ipynb"],72    "Kotlin" : [".kt"],73    "Lisp" : [".lisp"],74    "Matlab" : [".m"],75    "PHP": [".php"],76    "Perl": [".pl"],77    "Python": [".py"],78    "R" : [".r"],79    "Ruby": [".rb"],80    "Rust": [".rs"],81    "SQL": [".sql"],82    "Shell": [".sh"],83    "Swift" : [".swift"],84    "TypeScript": [".ts"],85}86```87## Licenses88Each example is also annotated with the license of the associated repository. There are in total 15 licenses:89```python90[91    'mit',92    'apache-2.0',93    'gpl-2.0',94    'gpl-3.0',95    'bsd-3-clause',96    'bsd-2-clause',97    'unlicense',98    'apacheagpl-3.0',99    'lgpl-3.0',100    'cc0-1.0',101    'epl-1.0',102    'lgpl-2.1',103    'mpl-2.0',104    'isc',105    'artistic-2.0'106 ]107```108## Dataset Statistics109The dataset is about ~ 18 TB uncompressed. We are currently working on processing it and applying further filtering.110## Dataset Creation111The dataset was created in two steps:1121. Files with the extensions given in the list above were retrieved from the GitHub dataset on BigQuery using the following query:113 114```sql115SELECT116  f.id, f.repo_name, f.path, content.copies, content.size, content.content, lic.license117FROM118  `bigquery-public-data.github_repos.files` AS f119JOIN120 121  `bigquery-public-data.github_repos.contents` as content122ON123  f.id = content.id124JOIN125  `bigquery-public-data.github_repos.licenses` AS lic126 127ON128  f.repo_name = lic.repo_name 129 130 131WHERE132    NOT content.binary133    134 135      136    AND (137 138        (f.path LIKE '%.py') OR (f.path LIKE '%.java')  OR (f.path LIKE '%.js') 139         OR (f.path LIKE '%.html') OR (f.path LIKE '%.lisp') OR (f.path LIKE '%.sh') 140         OR (f.path LIKE '%.r') OR (f.path LIKE '%.pl') OR (f.path LIKE '%.css')141         OR (f.path LIKE '%.sql') OR (f.path LIKE '%.c') OR (f.path LIKE '%.cpp') 142         OR (f.path LIKE '%.ts') OR (f.path LIKE '%.cs') OR (f.path LIKE '%.go') 143         OR (f.path LIKE '%.rs') OR (f.path LIKE '%.swift') OR (f.path LIKE '%.php') 144         OR (f.path LIKE '%.dart') OR (f.path LIKE '%.kt') OR (f.path LIKE '%.m') 145         OR (f.path LIKE '%.rb') OR (f.path LIKE '%.ipynb')146        147    )148     149    -- make sure we dont go above 1 megabyte  150    AND (content.size BETWEEN 1024 AND 1000000)151 152 153 154```1552. Currently, our CodedotAI team is working on adding additional filters and cleaning this dataset.156### Personal and Sensitive Information157 158Since this data was collected from public repositories, there exists potential for personal and sensitive information to be included in the data through developers accidentally or on purpose uploading their secret keys, passwords, API keys, emails, etc.159 160## Considerations for Using the Data161 162### Social Impact of Dataset163 164The paper ["Evaluating Large Language Models Trained on Code"](https://arxiv.org/abs/2107.03374) from OpenAI has a good discussion on what the impact of a large language model trained on code could be. Therefore, some parts of their discussion are highlighted here as it pertains to this dataset and models that may be trained from it. **As well as some differences in views from the paper, particularly around legal implications**.165 1661. **Over-reliance:** A language model trained on large datasets such as this one for the task of autogenerating code may generate plausible solutions that may appear correct, but are not necessarily the correct solution. Not properly evaluating the generated code may cause have negative consequences such as the introduction of bugs, or the introduction of security vulnerabilities. Therefore, it is important that users are aware of the limitations and potential negative consequences of using a language model trained on this dataset.1672. **Economic and labor market impacts:** Large language models trained on large code datasets such as this one that are capable of generating high-quality code have the potential to automate part of the software development process. This may negatively impact software developers. However, as discussed in the paper, as shown in the Summary Report of software developers from [O*NET OnLine](https://www.onetonline.org/link/summary/15-1252.00), developers don't just write software.1683. **Security implications:** No filtering or checking of vulnerabilities or buggy code was performed. This means that the dataset may contain code that may be malicious or contain vulnerabilities. Therefore, any model trained on this dataset may generate vulnerable, buggy, or malicious code. In safety-critical software, this could lead to software that may work improperly and could result in serious consequences depending on the software. Additionally, a model trained on this dataset may be used to generate malicious code on purpose in order to perform ransomware or other such attacks.1694. **Legal implications:** No filtering was performed on licensed code. This means that the dataset may contain restrictive licensed code. As discussed in the paper, public Github repositories may fall under "fair use." However, there have been little to no previous cases of such usages of licensed publicly available code. Therefore, any model trained on this dataset may be required to obey license terms that align with the software it was trained on such as GPL-3.0, which is why we purposefully put this dataset under the GPL-3.0 license. It is unclear the legal ramifications of using a language model trained on this dataset.170 171 172### v1.0173- The query was executed on _February 1, 2022, 12:15:59 AM EST_174 175## Acknowledgements176This project would not have been possible without compute generously provided by Google through the [TPU Research Cloud](https://sites.research.google/trc/about/). We would also like to thank [Dr. Razvan Bunescu](https://webpages.charlotte.edu/rbunescu/) and [The College of Computing and Informatics at UNC Charlotte](https://cci.charlotte.edu/) for their generous contributions to this project, specifically in funding the BigQuery and Google Cloud Storage costs. We would also like to thank the [codeparrot team at Hugging face](https://huggingface.co/codeparrot) for open sourcing their documentation on [github-code](https://huggingface.co/datasets/codeparrot/github-code) which we used for the readme in this dataset. For another similar dataset to this please check github-code!