Team Ai
Datasetpublic

Aditya78b/codeparrot-java-all

GitHub Code Dataset Dataset Description The GitHub Code dataset consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in 1TB of data. The dataset was created from the public GitHub dataset on Google BiqQuery. How to use it The GitHub Code dataset is a very large dataset so for most use cases it is recommended to make use of the streaming API of datasets. You can load and iterate through the dataset with the… See the full description on the dataset page: https://huggingface.co/datasets/Aditya78b/codeparrot-java-all.

sourceHugging Faceotherupdated 3y agoView on Hugging Face
0likes31downloads
4 commits on main
20dfe5d3y ago

Upload train-00000-of-01126.parquet

Aditya78b
9f0c4db3y ago

Delete data/train-00000-of-01126.parquet

Aditya78b
4d856833y ago

Upload 7 files

Aditya78b
96a44773y ago

initial commit

Aditya78b