Team Ai
Datasetpublic

ajibawa-2023/Java-Code-Large

Java-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Java-Code-Large.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
34likes502downloads
10 commits on main
be47bf98mo ago

Upload 10 files

ajibawa-2023
0cd2e3d8mo ago

Delete Java

ajibawa-2023
9ad2f7c8mo ago

Create Java

ajibawa-2023
1014ed98mo ago

Upload 10 files

ajibawa-2023
f19de4b8mo ago

Upload 20 files

ajibawa-2023
7c69dba8mo ago

Upload 13 files

ajibawa-2023
452746f8mo ago

Upload 18 files

ajibawa-2023
d6494728mo ago

Update README.md

ajibawa-2023
915a0df8mo ago

Update README.md

ajibawa-2023
c5d6bcb8mo ago

initial commit

ajibawa-2023