datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
starcoder2data-extras
StarCoder2 Extras
This is the dataset of extra sources (besides Stack v2 code data) used to train the StarCoder2 family of models. It contains the following subsets:
Kaggle (kaggle): Kaggle notebooks from Meta-Kaggle-Code dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script.
StackOverflow (stackoverflow): stackoverflow… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/starcoder2data-extras.stack-v2-starcoder2-3bstarcoder2data-extras
StarCoder2 Extras
This is the dataset of extra sources (besides Stack v2 code data) used to train the StarCoder2 family of models. It contains the following subsets:
Kaggle (kaggle): Kaggle notebooks from Meta-Kaggle-Code dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script.
StackOverflow (stackoverflow): stackoverflow… See the full description on the dataset page: https://huggingface.co/datasets/animeshjoshi0086/starcoder2data-extras.starcoder2-instruct-assetsstarcoder2-documentation
Dataset Card
This dataset is the code documenation dataset used in StarCoder2 pre-training, and it is also part of the-stack-v2-train-extras descried in the paper.
Dataset Details
Overview
This dataset comprises a comprehensive collection of crawled documentation and code-related resources sourced from various package manager platforms and programming language documentation sites. It focuses on popular libraries, free programming books, and other relevant… See the full description on the dataset page: https://huggingface.co/datasets/SivilTaram/starcoder2-documentation.details_cognitivecomputations__dolphincoder-starcoder2-7bstarcoder2-evaluationThis is not a Hugging Face dataset. We are just using it as a Git repository. Click on Files and versions above to see what's here.
dolma-blend-starcoder2details_bigcode__starcoder2-7b
Dataset Card for Evaluation run of bigcode/starcoder2-7b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-7b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigcode__starcoder2-7b.jailbreaks_dataset_with_perplexity_bigcode_starcoder2-3b_bigcode_starcoder2-7bbigcode__starcoder2-3b-details
Dataset Card for Evaluation run of bigcode/starcoder2-3b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-3b-details.details_bigcode__starcoder2-3b
Dataset Card for Evaluation run of bigcode/starcoder2-3b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-3b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigcode__starcoder2-3b.bigcode__starcoder2-15b-details
Dataset Card for Evaluation run of bigcode/starcoder2-15b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-15b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-15b-details.bigcode__starcoder2-7b-details
Dataset Card for Evaluation run of bigcode/starcoder2-7b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-7b-details.details_bigcode__starcoder2-15b
Dataset Card for Evaluation run of bigcode/starcoder2-15b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-15b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bigcode__starcoder2-15b.starcoder2-selfinstructreward-bench-starcoder2-7b-yes-nothe-stack-v2-dedup-python-starcoder2-3bBlenderCAD2-Ollama-Starcoder2-7blocal-code-arena-starcoder2_3b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder2 3B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the next-generation StarCoder2 3B base foundational model.
This specific partition documents the behavioral dynamics of modern raw foundational weights inside automated conversational pipelines, highlighting the persistent… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder2_3b.local-code-arena-starcoder2_7b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder2 7B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the next-generation StarCoder2 7B base foundational model.
This specific partition documents the behavioral dynamics of modern, mid-tier raw foundational weights inside automated conversational evaluation workflows, defining the… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder2_7b.local-code-arena-starcoder2_15b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder2 15B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the flagship StarCoder2 15B base foundational model.
This specific partition documents the final limits of scaling raw, unaligned foundational weights inside conversational evaluation loops, establishing an absolute baseline for… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder2_15b.ontocord__starcoder2_3b-AutoRedteam-details
Dataset Card for Evaluation run of ontocord/starcoder2_3b-AutoRedteam
Dataset automatically created during the evaluation run of model ontocord/starcoder2_3b-AutoRedteam
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2_3b-AutoRedteam-details.ontocord__starcoder2-29b-ls-details
Dataset Card for Evaluation run of ontocord/starcoder2-29b-ls
Dataset automatically created during the evaluation run of model ontocord/starcoder2-29b-ls
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2-29b-ls-details.instruction_response_starcoder2-15bstarcoder2-3b-js
