Team Ai
Datasetpublic

HuggingFaceCode/stack-v3-train

🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.

sourceHugging Faceodc-byupdated 15d agoView on Hugging Face
386likes160kdownloads
banner.png4 linesDownload Raw Back to assets
1version https://git-lfs.github.com/spec/v12oid sha256:cdeb08746abf0928498195dd484a7e2399f87019a3912b373c6bc73738bb6f2f3size 15673044 
HuggingFaceCode/stack-v3-train · Team Ai