Team Ai
Datasetpublic

codeparrot/github-code-clean

The GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in almost 1TB of text data.

sourceHugging Faceapache-2.0updated 4y agoView on Hugging Face
145likes31kdownloads
README.md9 linesDownload Raw Back to root
1---2license: apache-2.03---4This is a cleaner version of [Github-code dataset](https://huggingface.co/datasets/codeparrot/github-code), we add the following filters:5* Average line length < 1006* Alpha numeric characters fraction > 0.257* Remove auto-generated files (keyword search)8 93.39M files are removed making up 2.94% of the dataset.