Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigcode /starcoder2data-extras StarCoder2 Extras This is the dataset of extra sources (besides Stack v2 code data) used to train the StarCoder2 family of models. It contains the following subsets: Kaggle (kaggle): Kaggle notebooks from Meta-Kaggle-Code dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script. StackOverflow (stackoverflow): stackoverflow… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/starcoder2data-extras.tabular10M<n<100M13 likes3.3k downloads2y agoHugging Face02jon-tow /starcoderdata-python-edu starcoderdata-python-edu StarCoder Training Dataset Cleaned and Scored Dataset Details Dataset Description This dataset is a filtered version of StarCoder Training Dataset that has been scored with the python-edu-scorer. Dataset Sources Repository: https://huggingface.co/collections/HuggingFaceTB/smollm-6695016cad7167254ce15966 Paper: SmolLM - blazingly fast and remarkably powerful Citation @misc{allal2024SmolLM, title={SmolLM… See the full description on the dataset page: https://huggingface.co/datasets/jon-tow/starcoderdata-python-edu.tabular10M<n<100M14 likes2k downloads2y agoHugging Face03animeshjoshi0086 /starcoder2data-extras StarCoder2 Extras This is the dataset of extra sources (besides Stack v2 code data) used to train the StarCoder2 family of models. It contains the following subsets: Kaggle (kaggle): Kaggle notebooks from Meta-Kaggle-Code dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script. StackOverflow (stackoverflow): stackoverflow… See the full description on the dataset page: https://huggingface.co/datasets/animeshjoshi0086/starcoder2data-extras.tabular10M<n<100M0 likes1.9k downloads6mo agoHugging Face04vikp /starcoder_labeled Dataset Card for "starcoder_labeled" Starcoder data, with several popular languages selected, short sequences filtered out, then labeled based on learning quality (educational value) and code quality. A good heuristic is to take anything with >.5 code quality and >.3 learning quality. But you may want to vary the thresholds by language, depending on your target task. tabular10M<n<100M2 likes867 downloads3y agoHugging Face05JanSchTech /starcoderdata-python-edu-lang-score Dataset Card for Starcoder Data with Python Education and Language Scores Dataset Summary The starcoderdata-python-edu-lang-score dataset contains the Python subset of the starcoderdata dataset. It augments the existing Python subset with features that assess the educational quality of code and classify the language of code comments. This dataset was created for high-quality Python education and language-based training, with a primary focus on facilitating models that can… See the full description on the dataset page: https://huggingface.co/datasets/JanSchTech/starcoderdata-python-edu-lang-score.tabular1M<n<10M2 likes734 downloads2y agoHugging Face06ytzi /starcoderdata-gpt2tabular10M<n<100M0 likes423 downloads3y agoHugging Face07skymizer /common_starcoder Common Starcoder dataset This dataset is generated from bigcode/starcoderdata. Total GPT2 Tokens: 4,649,163,171 Generation Process We filtered the original dataset with common language: C, Cpp, Java, Python and JSON. We removed some columns for mixing up with other dataset: "id", "max_stars_repo_path", "max_stars_repo_name" After removing the irrelevant fields, we shuffle the dataset with random seed=42. We filtered the data on "max_stars_count" > 300 and shuffle again.… See the full description on the dataset page: https://huggingface.co/datasets/skymizer/common_starcoder.tabular1M<n<10M1 likes288 downloads2y agoHugging Face08Sam-Shin /starcoder Starcoder Dataset (The Stack - Sub-sampled) This dataset is derived from the "Starcoder" version of The Stack, a 6.4 TB dataset of permissively licensed source code in 384 programming languages. This repository contains the data organized into subsets, one for each programming language or data type. How to Use You can load any language-specific subset of the data using the datasets library. You must specify the name parameter with the desired language. For example… See the full description on the dataset page: https://huggingface.co/datasets/Sam-Shin/starcoder.tabular100M<n<1B0 likes191 downloads11mo agoHugging Face09lparkourer10 /starcoder-python5b5b gpt2 tokens tabulartext-generation1M<n<10M0 likes166 downloads2y agoHugging Face10nuprl /stack-dedup-python-testgen-starcoder-filter-v2 MultiPL-T Python Sources Citation If you use this dataset we request that you cite our work: @misc{cassano:multipl-t, title={Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs}, author={Federico Cassano and John Gouwar and Francesca Lucchetti and Claire Schlesinger and Anders Freeman and Carolyn Jane Anderson and Molly Q Feldman and Michael Greenberg and Abhinav Jangda and Arjun Guha}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/stack-dedup-python-testgen-starcoder-filter-v2.tabular100K<n<1M7 likes161 downloads3y agoHugging Face11malaysia-ai /starcoderdata-sampletabular100K<n<1M0 likes150 downloads3y agoHugging Face12open-llm-leaderboard /bigcode__starcoder2-3b-detailsgated Dataset Card for Evaluation run of bigcode/starcoder2-3b Dataset automatically created during the evaluation run of model bigcode/starcoder2-3b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-3b-details.tabular10K<n<100K0 likes95 downloads2y agoHugging Face13open-llm-leaderboard /bigcode__starcoder2-15b-detailsgated Dataset Card for Evaluation run of bigcode/starcoder2-15b Dataset automatically created during the evaluation run of model bigcode/starcoder2-15b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-15b-details.tabular10K<n<100K0 likes69 downloads2y agoHugging Face14open-llm-leaderboard /bigcode__starcoder2-7b-detailsgated Dataset Card for Evaluation run of bigcode/starcoder2-7b Dataset automatically created during the evaluation run of model bigcode/starcoder2-7b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-7b-details.tabular10K<n<100K0 likes63 downloads2y agoHugging Face15mdonigian /starcoder-curated StarCoderData Curated A curated subset of StarCoderData optimised for training a 500M parameter model focused on structured data output (JSON generation, function calling, schema compliance). Dataset Summary Total code files: 5,203,508 Total tokens: 3.9B (target: 3.5B) Classifier-scored files: 1,553,596 (1.7B tokens) Non-classified files: 3,649,912 (2.2B tokens) — filtered by heuristics, not the classifier Source: bigcode/starcoderdata Classifier:… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/starcoder-curated.imagetext-generation1M<n<10M0 likes51 downloads8mo agoHugging Face16loubnabnl /starcoderdata_py_smol Dataset Card for "starcoderdata_py_smol" More Information needed tabular100K<n<1M1 likes33 downloads3y agoHugging Face17Hietan /starcoderdata_100star_py_stripped_v1tabular100K<n<1M0 likes33 downloads1y agoHugging Face18Ayush-Singh /reward-bench-starcoder2-7b-yes-notabularn<1K0 likes18 downloads2y agoHugging Face19nuprl /stack-dedup-python-testgen-starcoder-filter-inferred-v2 Dataset Card for "stack-dedup-python-testgen-starcoder-filter-inferred-v2" More Information needed tabular100K<n<1M1 likes17 downloads3y agoHugging Face20zhaospei /starcoder_java_finetunetabular1K<n<10K0 likes12 downloads2y agoHugging Face21open-llm-leaderboard /ontocord__starcoder2_3b-AutoRedteam-detailsgated Dataset Card for Evaluation run of ontocord/starcoder2_3b-AutoRedteam Dataset automatically created during the evaluation run of model ontocord/starcoder2_3b-AutoRedteam The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2_3b-AutoRedteam-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face22open-llm-leaderboard /ontocord__starcoder2-29b-ls-detailsgated Dataset Card for Evaluation run of ontocord/starcoder2-29b-ls Dataset automatically created during the evaluation run of model ontocord/starcoder2-29b-ls The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2-29b-ls-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face23nuprl /stack-dedup-python-testgen-starcoder-filter-v2-dedupgated MultiPL-T Python Sources Citation If you use this dataset we request that you cite our work: @misc{cassano:multipl-t, title={Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs}, author={Federico Cassano and John Gouwar and Francesca Lucchetti and Claire Schlesinger and Anders Freeman and Carolyn Jane Anderson and Molly Q Feldman and Michael Greenberg and Abhinav Jangda and Arjun Guha}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/stack-dedup-python-testgen-starcoder-filter-v2-dedup.tabular100K<n<1M1 likes7 downloads3y agoHugging Face24rajlohith2 /starcoder_filteredtabular1K<n<10K0 likes7 downloads2y agoHugging Face25echodrift /starcoder-finetune-javatabular10K<n<100K0 likes6 downloads2y agoHugging Face26zhaospei /starcoder_java_refinetabular10K<n<100K0 likes4 downloads2y agoHugging Face27Hietan /starcoderdata_100star_py_annotated_v1tabular100K<n<1M0 likes4 downloads1y agoHugging Face28zhaospei /gemma-starcoder-baseline-soliditytabular1K<n<10K1 likes3 downloads2y agoHugging Face29zhaospei /starcoder_3b_baseline_soliditytabular1K<n<10K1 likes3 downloads2y agoHugging Face30zhaospei /starcoder_java_baselinetabular1K<n<10K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.