datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_contests_slim_jsoncode_contests_dp_datasetcodecomplex
CodeComplex Dataset
Dataset Description
CodeComplex consists of 4,200 Java codes submitted to programming competitions by human programmers and their complexity labels annotated by a group of algorithm experts.
How to use it
You can load and iterate through the dataset with the following two lines of code:
from datasets import load_dataset
ds = load_dataset("codeparrot/codecomplex", split="train")
print(next(iter(ds)))
Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/codecomplex.code-code-DefectDetection
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Defect-detection in Semeru
CodeXGLUE -- Defect Detection
Task Definition
Given a source code, the task is to identify whether it is an insecure code that may attack software systems, such as resource leaks, use-after-free vulnerabilities and DoS attack. We treat the task as… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-DefectDetection.code_contestscode-comprehension
Dataset Card for Code Comprehension
Collection of code understanding questions.
Dataset Details
Dataset Description
These examples fall into 2 categories:
"cloze": fill in the hole to produce the specified outcome;
"eval": given a snippet of python code, determine the outcome.
Some questions are very easy, some are much more challenging.
Most (if not all) of these questions should be relatively straightforward
for an experienced programmer, even without a… See the full description on the dataset page: https://huggingface.co/datasets/imbue/code-comprehension.CodeCrashCodeContests_trainWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format:
input = "5\n1 2 3 4 5\n"
output = "15"
CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like
assert sum_function([1, 2, 3, 4, 5]) == 15
In this project, we have converted the the functional format to the Stdio format to achieve consistency.
Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests_train.code-code-galeras-code-completion-from-docstring-3k-dedupedcritic-sft-cwm-only
critic-sft-cwm-only
The CWM-only SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-CWM-only, the CWM-only arm of the corpus ablation in Table 3.
Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it under the paper's high-level prompt. The main corpus, critic-sft-cwm-qwen, is this set plus 1,915 records from Qwen3-Next trajectories.… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only.code-contract-repaircritic-sft-cwm-only-detailed-prompt
critic-sft-cwm-only-detailed-prompt
The detailed-prompt SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Detailed-Prompt, the comparison arm of the prompt ablation in Table 4.
Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The difference from critic-sft-cwm-only is the teacher prompt. Here the teacher used the detailed prompt… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only-detailed-prompt.critic-sft-cwm-qwen
critic-sft-cwm-qwen
The main SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains both Qwen3-8B-Critic-SFT and Qwen3-4B-Critic-SFT.
Each record is one critique point: a coding agent's trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The teacher was prompted with the paper's high-level prompt, which asks for error detection and one or two sentences of guidance and forbids code and commands in the… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-qwen.CodeContestsWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format:
input = "5\n1 2 3 4 5\n"
output = "15"
CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like
assert sum_function([1, 2, 3, 4, 5]) == 15
In this project, we have converted the the functional format to the Stdio format to achieve consistency.
Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests.code-code-MethodGeneration
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Method-Generation/dataset/codexglue_method_generation in Semeru
CodeXGLUE -- Method Generation
Here is the introduction and pipeline for method generation task.
Task Definition
Method generation is the prediction of a method body implementation conditioned on a signature, a… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-MethodGeneration.code-contract-repair
APIContractRepair
APIContractRepair is a provenance-tracked instruction-tuning dataset for software engineers and code-model researchers who need contract-faithful, minimal repairs with tests that distinguish a broken implementation from its fix. Magicoder-OSS-Instruct-75K supplies real function-identifier seeds, but it does not provide these documented contracts, deliberately buggy implementations, minimal corrected implementations, or paired regression tests. This release… See the full description on the dataset page: https://huggingface.co/datasets/skonml/code-contract-repair.code-cpt-corpus
Dataset Card for code-cpt-corpus
Dataset Summary
ဒီ dataset က code-cpt-corpus အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
data/train.jsonl
Licensing Information
ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/code-cpt-corpus.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/codecainecowboy/GLM-5.1-Reasoning-1M-Cleaned.codecomplex
CodeComplex Dataset
Dataset Description
CodeComplex consists of 4,200 Java codes submitted to programming competitions by human programmers and their complexity labels annotated by a group of algorithm experts.
How to use it
You can load and iterate through the dataset with the following two lines of code:
from datasets import load_dataset
ds = load_dataset("codeparrot/codecomplex", split="train")
print(next(iter(ds)))
Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/HlebYakhnitski/codecomplex.critic-sft-qwen-only
critic-sft-qwen-only
The Qwen-only SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Qwen-only and Qwen3-4B-Critic-SFT-Qwen-only, the Qwen-only arms of the corpus ablation in Table 3.
Each record is one critique point: a Qwen3-Next-80B-A3B-Instruct trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it under the paper's high-level prompt. The main corpus, critic-sft-cwm-qwen, is… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-qwen-only.PRM_1541i
Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper.
PRM_1541i
1541 preference pairs for training a critic over coding-agent trajectories. Each example is a
multi-turn agent transcript paired with two candidate… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/PRM_1541i.base_dataset_07_nov_23This dataset is 10% repo sampled dataset for selected languages. We applied a repo sample rate of 10%. e.g. if sample rate is 10% then we take 10% of all repos for a
given language but include all files inside the repo.
This was generated using our codecomplete/training/completions/datagen
./launch.sh \
--dataset-name bigcode/starcoderdata \
--subset c,cpp,go,java,javascript,typescript,python,ruby,scala,sql \
--sample-rate 0.01 \
--hf-token <HF_TOKEN> \
--output-dir… See the full description on the dataset page: https://huggingface.co/datasets/codecompletedeployment/base_dataset_07_nov_23.base_datasetThis dataset is 10% repo sampled dataset for selected languages. We applied a repo sample rate of 10%. e.g. if sample rate is 10% then we take 10% of all repos for a
given language but include all files inside the repo.
This was generated using our codecomplete/training/completions/datagen
./launch.sh \
--dataset-name bigcode/starcoderdata \
--subset c,cpp,go,java,javascript,typescript,python,ruby,scala,sql \
--sample-rate 0.01 \
--hf-token <HF_TOKEN> \
--output-dir… See the full description on the dataset page: https://huggingface.co/datasets/codecomplete/base_dataset.Code-Code-CloneDetection-POJ104
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Clone-detection-POJ-104 in Semeru
CodeXGLUE -- Clone Detection (POJ-104)
Task Definition
Given a code and a collection of candidates as the input, the task is to return Top K codes with the same semantic. Models are evaluated by MAP@R score. MAP@R is defined as the mean of… See the full description on the dataset page: https://huggingface.co/datasets/semeru/Code-Code-CloneDetection-POJ104.Code-code-galeras-prompting-3k-treatment-1codeconv-fortran-to-rustcodecCode-code-galeras-prompting-3k-treatment-2audio_codecCodeComp2
