Team Ai
Datasetpublic

HexQuant/Code-Contests-Plus

CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset built upon CodeContests. It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions. Highlights High… See the full description on the dataset page: https://huggingface.co/datasets/HexQuant/Code-Contests-Plus.

sourceHugging Facecc-by-4.0updated 10mo agoView on Hugging Face
1likes727downloads
README.md134 linesDownload Raw Back to root
1---2license: cc-by-4.03size_categories:4- 10K<n<100K5tags:6- code7task_categories:8- other9configs:10- config_name: default11  data_files:12  - split: train13    path: part-*14- config_name: 1x15  data_files:16  - split: train17    path: ccplus_1x/*18- config_name: 2x19  data_files:20  - split: train21    path: ccplus_2x/*22- config_name: 3x23  data_files:24  - split: train25    path: ccplus_3x/*26- config_name: 4x27  data_files:28  - split: train29    path: ccplus_4x/*30- config_name: 5x31  data_files:32  - split: train33    path: ccplus_5x/*34---35 36<div align="center">37  <h1>CodeContests<sup>+</sup>: A Competitive Programming Dataset with High-Quality Test Cases</h1>38</div>39 40<div align="center" style="line-height: 1;">41  <a href="https://arxiv.org/abs/2506.05817" target="_blank" style="margin: 2px;">42    <img alt="2506.05817" src="https://img.shields.io/badge/arXiv-2506.05817-red?logo=arxiv&logoColor=white" style="display: inline-block; vertical-align: middle;"/>43  </a>44  <a href="https://huggingface.co/datasets/ByteDance-Seed" target="_blank" style="margin: 2px;">45    <img alt="Hugging Face" src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-ByteDance%20Seed-536af5" style="display: inline-block; vertical-align: middle;"/>46  </a>47  <a href="https://huggingface.co/datasets/ByteDance-Seed/Code-Contests-Plus/blob/main/LICENSE" style="margin: 2px;">48    <img alt="Dataset License" src="https://img.shields.io/badge/Dataset_License-CC--BY--4.0-f5de53?&color=f5de53" style="display: inline-block; vertical-align: middle;"/>49  </a>50  <a href="https://github.com/bytedance/SandboxFusion/blob/main/LICENSE" style="margin: 2px;">51    <img alt="Sandbox License" src="https://img.shields.io/badge/Sandbox_License-Apache--2.0-f5de53?&color=f5de53" style="display: inline-block; vertical-align: middle;"/>52  </a>53</div>54 55## Introduction56 57CodeContests<sup>+</sup> is a competitive programming problem dataset built upon [CodeContests](https://huggingface.co/datasets/deepmind/code_contests). It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions.58 59## Highlights60 61**High Quality Test Cases:** We developed a Generator-Validator Agent System that can construct high-quality test cases for each problem. In addition to random test cases, it also generates special test cases tailored to the problem's characteristics and various corner cases, aiming to cover as many potential errors as possible. Furthermore, the correctness of these test cases is verified by an independent test case validator to ensure they comply with the problem constraints.62 63**Test Case Generators:** We provide a test case generator for each problem, along with the commands to run it. These commands can be run multiple times to produce an infinite number of test cases. This allows users to understand the specific characteristics of all test cases clearly and enables them to use these generators to create as many additional test cases as they need.64 65**Flexible Number of Test Cases:** Additionally, we also provide pre-generated test cases, available in five versions: 1x, 2x, ..., 5x. The number of test cases in these versions increases sequentially, so the computational resources required to run them will also increase. This allows users to strike a balance between computational cost and coverage according to their needs.66 67**Test Case Validators:** Competitive programming problems usually specify many constraints on the input data itself, including data ranges, format requirements, data structure requirements, and so on. Therefore, constructing fully valid test cases is not an easy task, and even professional problem setters can easily make mistakes. For each problem, we provide a test case validator that strictly checks whether the test case input satisfies all constraints outlined in the problem description, to ensure the validity of the test cases as much as possible.68 69**Output Checkers for Multiple Answer Problems:** In programming competitions, problems with multiple valid solutions are very common. This means that the same input can correspond to several valid outputs. Therefore, correctness cannot be determined simply by comparing the program's output with a single, pre-defined correct answer. For this reason, we provide custom output checkers for all such problems to verify the correctness of the output.70 71**Rigorous Evaluation:** To rigorously evaluate the quality of these test cases, we assessed their accuracy using a large number of solutions. For each problem, we used 100 correct solutions and 100 incorrect solutions to determine if the test cases could correctly distinguish between correct and incorrect submissions. We have recorded the evaluation results, including True Positive Rate (TPR) and True Negative Rate (TNR), in the dataset. Additionally, based on these results, we selected a high-quality subset from the full dataset, named [CodeContests<sup>+</sup>Verified](https://huggingface.co/datasets/ByteDance-Seed/Code-Contests-Plus-Verified), in which the TPR and TNR for each problem are both above 0.9. Users can apply their own filtering if they require a looser or stricter threshold72 73## Quickstart74 75Load dataset without test cases:76```python77from datasets import load_dataset78 79# Login using e.g. `huggingface-cli login` to access this dataset80ds = load_dataset("ByteDance-Seed/Code-Contests-Plus", "default")81```82 83Load dataset with `1x` test cases:84```python85from datasets import load_dataset86 87# Login using e.g. `huggingface-cli login` to access this dataset88ds = load_dataset("ByteDance-Seed/Code-Contests-Plus", "1x")89```90 91This dataset has 6 subsets, namely `default`, `1x`, `2x`, `3x`, `4x`, and `5x`. The problems in these subsets are the same. The only difference is the number of test cases. The table below presents the average number of test cases per problem in each subset. 92 93| Subset            | default | 1x | 2x | 3x | 4x | 5x |94|-------------------|---------|----|----|----|----|----|95| Avg. # test cases | 0       | 25 | 44 | 62 | 80 | 98 |96 97## Usage98 99We recommend using CodeContests<sup>+</sup> with [SandboxFusion](https://github.com/bytedance/SandboxFusion). SandboxFusion supports the automatic evaluation on 10+ open-source datasets, including CodeContest<sup>+</sup>, LiveCodeBench, HumanEval, MBPP, MHPP, and 20+ programming languages, including C++, Python (GPU supported), C#, Go, Java, NodeJS, Typescript, Kotlin, Rust, Bash, PHP, and even Verilog.100 101## Evaluation Result102 103![Fig](https://huggingface.co/datasets/ByteDance-Seed/Code-Contests-Plus/resolve/main/result.png)104 105We present the histogram of the TPR and TNR of problems from (a) CodeContests and (b)CodeContests<sup>+</sup> above. For more details of our evaluation, please refer to our [paper](https://arxiv.org/abs/2506.05817).106 107## License108 109This project is licensed under CC-BY-4.0. See the [LICENSE file](https://huggingface.co/datasets/ByteDance-Seed/Code-Contests-Plus/blob/main/LICENSE) for details.110 111## Citation112 113```114@inproceedings{wang-etal-2025-codecontests,115    title = "{C}ode{C}ontests+: High-Quality Test Case Generation for Competitive Programming",116    author = "Wang, Zihan  and117      Liu, Siyao  and118      Sun, Yang  and119      Ding, Ming  and120      Li, Hongyan",121    editor = "Christodoulopoulos, Christos  and122      Chakraborty, Tanmoy  and123      Rose, Carolyn  and124      Peng, Violet",125    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025",126    month = nov,127    year = "2025",128    address = "Suzhou, China",129    publisher = "Association for Computational Linguistics",130    url = "https://aclanthology.org/2025.findings-emnlp.299/",131    pages = "5576--5600",132    ISBN = "979-8-89176-335-7"133}134```