datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
starcoder2data-extras
StarCoder2 Extras
This is the dataset of extra sources (besides Stack v2 code data) used to train the StarCoder2 family of models. It contains the following subsets:
Kaggle (kaggle): Kaggle notebooks from Meta-Kaggle-Code dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script.
StackOverflow (stackoverflow): stackoverflow… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/starcoder2data-extras.starcoderdata-python-edu
starcoderdata-python-edu
StarCoder Training Dataset Cleaned and Scored
Dataset Details
Dataset Description
This dataset is a filtered version of StarCoder Training Dataset
that has been scored with the python-edu-scorer.
Dataset Sources
Repository: https://huggingface.co/collections/HuggingFaceTB/smollm-6695016cad7167254ce15966
Paper: SmolLM - blazingly fast and remarkably powerful
Citation
@misc{allal2024SmolLM,
title={SmolLM… See the full description on the dataset page: https://huggingface.co/datasets/jon-tow/starcoderdata-python-edu.starcoder2data-extras
StarCoder2 Extras
This is the dataset of extra sources (besides Stack v2 code data) used to train the StarCoder2 family of models. It contains the following subsets:
Kaggle (kaggle): Kaggle notebooks from Meta-Kaggle-Code dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script.
StackOverflow (stackoverflow): stackoverflow… See the full description on the dataset page: https://huggingface.co/datasets/animeshjoshi0086/starcoder2data-extras.starcoder_labeled
Dataset Card for "starcoder_labeled"
Starcoder data, with several popular languages selected, short sequences filtered out, then labeled based on learning quality (educational value) and code quality.
A good heuristic is to take anything with >.5 code quality and >.3 learning quality. But you may want to vary the thresholds by language, depending on your target task.
starcoderdata-python-edu-lang-score
Dataset Card for Starcoder Data with Python Education and Language Scores
Dataset Summary
The starcoderdata-python-edu-lang-score dataset contains the Python subset of the starcoderdata dataset. It augments the existing Python subset with features that assess the educational quality of code and classify the language of code comments. This dataset was created for high-quality Python education and language-based training, with a primary focus on facilitating models that can… See the full description on the dataset page: https://huggingface.co/datasets/JanSchTech/starcoderdata-python-edu-lang-score.starcoderdata-gpt2common_starcoder
Common Starcoder dataset
This dataset is generated from bigcode/starcoderdata.
Total GPT2 Tokens: 4,649,163,171
Generation Process
We filtered the original dataset with common language: C, Cpp, Java, Python and JSON.
We removed some columns for mixing up with other dataset: "id", "max_stars_repo_path", "max_stars_repo_name"
After removing the irrelevant fields, we shuffle the dataset with random seed=42.
We filtered the data on "max_stars_count" > 300 and shuffle again.… See the full description on the dataset page: https://huggingface.co/datasets/skymizer/common_starcoder.starcoder
Starcoder Dataset (The Stack - Sub-sampled)
This dataset is derived from the "Starcoder" version of The Stack, a 6.4 TB dataset of permissively licensed source code in 384 programming languages.
This repository contains the data organized into subsets, one for each programming language or data type.
How to Use
You can load any language-specific subset of the data using the datasets library. You must specify the name parameter with the desired language.
For example… See the full description on the dataset page: https://huggingface.co/datasets/Sam-Shin/starcoder.starcoder-python5b5b gpt2 tokens
stack-dedup-python-testgen-starcoder-filter-v2
MultiPL-T Python Sources
Citation
If you use this dataset we request that you cite our work:
@misc{cassano:multipl-t,
title={Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs},
author={Federico Cassano and John Gouwar and Francesca Lucchetti and Claire Schlesinger and Anders Freeman and Carolyn Jane Anderson and Molly Q Feldman and Michael Greenberg and Abhinav Jangda and Arjun Guha},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/stack-dedup-python-testgen-starcoder-filter-v2.starcoderdata-samplebigcode__starcoder2-3b-details
Dataset Card for Evaluation run of bigcode/starcoder2-3b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-3b-details.bigcode__starcoder2-15b-details
Dataset Card for Evaluation run of bigcode/starcoder2-15b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-15b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-15b-details.bigcode__starcoder2-7b-details
Dataset Card for Evaluation run of bigcode/starcoder2-7b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-7b-details.starcoder-curated
StarCoderData Curated
A curated subset of StarCoderData
optimised for training a 500M parameter model focused on structured data output
(JSON generation, function calling, schema compliance).
Dataset Summary
Total code files: 5,203,508
Total tokens: 3.9B (target: 3.5B)
Classifier-scored files: 1,553,596 (1.7B tokens)
Non-classified files: 3,649,912 (2.2B tokens) — filtered by heuristics, not the classifier
Source: bigcode/starcoderdata
Classifier:… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/starcoder-curated.starcoderdata_py_smol
Dataset Card for "starcoderdata_py_smol"
More Information needed
starcoderdata_100star_py_stripped_v1reward-bench-starcoder2-7b-yes-nostack-dedup-python-testgen-starcoder-filter-inferred-v2
Dataset Card for "stack-dedup-python-testgen-starcoder-filter-inferred-v2"
More Information needed
starcoder_java_finetuneontocord__starcoder2_3b-AutoRedteam-details
Dataset Card for Evaluation run of ontocord/starcoder2_3b-AutoRedteam
Dataset automatically created during the evaluation run of model ontocord/starcoder2_3b-AutoRedteam
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2_3b-AutoRedteam-details.ontocord__starcoder2-29b-ls-details
Dataset Card for Evaluation run of ontocord/starcoder2-29b-ls
Dataset automatically created during the evaluation run of model ontocord/starcoder2-29b-ls
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2-29b-ls-details.stack-dedup-python-testgen-starcoder-filter-v2-dedup
MultiPL-T Python Sources
Citation
If you use this dataset we request that you cite our work:
@misc{cassano:multipl-t,
title={Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs},
author={Federico Cassano and John Gouwar and Francesca Lucchetti and Claire Schlesinger and Anders Freeman and Carolyn Jane Anderson and Molly Q Feldman and Michael Greenberg and Abhinav Jangda and Arjun Guha},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/stack-dedup-python-testgen-starcoder-filter-v2-dedup.starcoder_filteredstarcoder-finetune-javastarcoder_java_refinestarcoderdata_100star_py_annotated_v1gemma-starcoder-baseline-soliditystarcoder_3b_baseline_soliditystarcoder_java_baseline
