Team Ai
Datasetpublic

Aditya78b/codeparrot-java-all

GitHub Code Dataset Dataset Description The GitHub Code dataset consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in 1TB of data. The dataset was created from the public GitHub dataset on Google BiqQuery. How to use it The GitHub Code dataset is a very large dataset so for most use cases it is recommended to make use of the streaming API of datasets. You can load and iterate through the dataset with the… See the full description on the dataset page: https://huggingface.co/datasets/Aditya78b/codeparrot-java-all.

sourceHugging Faceotherupdated 3y agoView on Hugging Face
0likes32downloads
github-code-stats-alpha.png4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:3c6020e90d24575c7d2e88d41a49754439320d0668e078a533852d3cb3a68d403size 1085624