Team Ai
Datasetpublic

Sam-Shin/starcoder

Starcoder Dataset (The Stack - Sub-sampled) This dataset is derived from the "Starcoder" version of The Stack, a 6.4 TB dataset of permissively licensed source code in 384 programming languages. This repository contains the data organized into subsets, one for each programming language or data type. How to Use You can load any language-specific subset of the data using the datasets library. You must specify the name parameter with the desired language. For example… See the full description on the dataset page: https://huggingface.co/datasets/Sam-Shin/starcoder.

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes191downloads
Dataset Card

Starcoder Dataset (The Stack - Sub-sampled)

This dataset is derived from the "Starcoder" version of The Stack, a 6.4 TB dataset of permissively licensed source code in 384 programming languages.

This repository contains the data organized into subsets, one for each programming language or data type.

How to Use

You can load any language-specific subset of the data using the datasets library. You must specify the name parameter with the desired language.

For example, to load the Python or Java subsets:

python
from datasets import load_dataset

# Load the Python subset
python_data = load_dataset("Sam-Shin/starcoder", name="python", split="train")

# Load the C++ subset
cpp_data = load_dataset("Sam-Shin/starcoder", name="cpp", split="train")

# Load the C# subset
csharp_data = load_dataset("Sam-Shin/starcoder", name="c-sharp", split="train")

print(python_data)