Team Ai
Datasetpublic

ammarnasr/the-stack-java-clean

Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.

sourceHugging Faceopenrailupdated 3y agoView on Hugging Face
13likes194downloads
bigcode-the-stack-dedup-train.jsonl4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:f84ab23610903da47a831f3c10e0e0a9d3d7f98c5b413bfa7368fb8ef9c05a3c3size 12346739504