Team Ai
Datasetpublic

Sam-Shin/starcoder

Starcoder Dataset (The Stack - Sub-sampled) This dataset is derived from the "Starcoder" version of The Stack, a 6.4 TB dataset of permissively licensed source code in 384 programming languages. This repository contains the data organized into subsets, one for each programming language or data type. How to Use You can load any language-specific subset of the data using the datasets library. You must specify the name parameter with the desired language. For example… See the full description on the dataset page: https://huggingface.co/datasets/Sam-Shin/starcoder.

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes191downloads
settings

This repository belongs to Sam-Shin on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namestarcoder
visibilitypublic
licencenot set
gatedno
ownerSam-Shin
Account settings
Sam-Shin/starcoder · Team Ai