Team Ai
Datasetpublic

XEUIPR/Java-Code-Large-text-only

Java-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.

sourceHugging Facemitupdated 3d agoView on Hugging Face
0likes3.8kdownloads
README.md26 linesDownload Raw Back to root
1---2license: mit3task_categories:4- text-generation5language:6- en7tags:8- code9- java10size_categories:11- 10M<n<100M12---13 14**Java-Code-Large**15 16Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than **15 million** java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis.17 18By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.19 20 21-----22 23 24Duplicated from [ajibawa-2023/Java-Code-Large](https://huggingface.co/datasets/ajibawa-2023/Java-Code-Large), where the only difference is that the "language" column containing only "java" is removed.25 26Dataset updated using huggingface_hub datasets on colab.