Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taoroalin /code_contests_slim_jsontabular1K<n<10K0 likes249 downloads2y agoHugging Face02hieunguyenminh /code_contests_dp_datasettabular1K<n<10K0 likes234 downloads2y agoHugging Face03codeparrot /codecomplex CodeComplex Dataset Dataset Description CodeComplex consists of 4,200 Java codes submitted to programming competitions by human programmers and their complexity labels annotated by a group of algorithm experts. How to use it You can load and iterate through the dataset with the following two lines of code: from datasets import load_dataset ds = load_dataset("codeparrot/codecomplex", split="train") print(next(iter(ds))) Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/codecomplex.texttext-generation1K<n<10K30 likes216 downloads4y agoHugging Face04semeru /code-code-DefectDetection Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Defect-detection in Semeru CodeXGLUE -- Defect Detection Task Definition Given a source code, the task is to identify whether it is an insecure code that may attack software systems, such as resource leaks, use-after-free vulnerabilities and DoS attack. We treat the task as… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-DefectDetection.tabular10K<n<100K2 likes128 downloads4y agoHugging Face05stamina /code_conteststext1M<n<10M0 likes121 downloads1y agoHugging Face06imbue /code-comprehension Dataset Card for Code Comprehension Collection of code understanding questions. Dataset Details Dataset Description These examples fall into 2 categories: "cloze": fill in the hole to produce the specified outcome; "eval": given a snippet of python code, determine the outcome. Some questions are very easy, some are much more challenging. Most (if not all) of these questions should be relatively straightforward for an experienced programmer, even without a… See the full description on the dataset page: https://huggingface.co/datasets/imbue/code-comprehension.text10K<n<100K21 likes98 downloads2y agoHugging Face07CUHK-ARISE /CodeCrashtext10K<n<100K4 likes76 downloads1y agoHugging Face08Gen-Verse /CodeContests_trainWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests_train.tabular1K<n<10K3 likes73 downloads1y agoHugging Face09semeru /code-code-galeras-code-completion-from-docstring-3k-dedupedtabular1K<n<10K3 likes71 downloads3y agoHugging Face10code-critic-model /critic-sft-cwm-only critic-sft-cwm-only The CWM-only SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-CWM-only, the CWM-only arm of the corpus ablation in Table 3. Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it under the paper's high-level prompt. The main corpus, critic-sft-cwm-qwen, is this set plus 1,915 records from Qwen3-Next trajectories.… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only.texttext-generation1K<n<10K0 likes69 downloads1mo agoHugging Face110xkamal7 /code-contract-repairtext1K<n<10K0 likes67 downloads2mo agoHugging Face12code-critic-model /critic-sft-cwm-only-detailed-prompt critic-sft-cwm-only-detailed-prompt The detailed-prompt SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Detailed-Prompt, the comparison arm of the prompt ablation in Table 4. Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The difference from critic-sft-cwm-only is the teacher prompt. Here the teacher used the detailed prompt… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only-detailed-prompt.texttext-generation1K<n<10K0 likes61 downloads1mo agoHugging Face13code-critic-model /critic-sft-cwm-qwen critic-sft-cwm-qwen The main SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains both Qwen3-8B-Critic-SFT and Qwen3-4B-Critic-SFT. Each record is one critique point: a coding agent's trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The teacher was prompted with the paper's high-level prompt, which asks for error detection and one or two sentences of guidance and forbids code and commands in the… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-qwen.texttext-generation1K<n<10K0 likes61 downloads1mo agoHugging Face14Gen-Verse /CodeContestsWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests.tabularn<1K1 likes60 downloads1y agoHugging Face15semeru /code-code-MethodGeneration Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Method-Generation/dataset/codexglue_method_generation in Semeru CodeXGLUE -- Method Generation Here is the introduction and pipeline for method generation task. Task Definition Method generation is the prediction of a method body implementation conditioned on a signature, a… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-MethodGeneration.text100K<n<1M1 likes54 downloads4y agoHugging Face16skonml /code-contract-repair APIContractRepair APIContractRepair is a provenance-tracked instruction-tuning dataset for software engineers and code-model researchers who need contract-faithful, minimal repairs with tests that distinguish a broken implementation from its fix. Magicoder-OSS-Instruct-75K supplies real function-identifier seeds, but it does not provide these documented contracts, deliberately buggy implementations, minimal corrected implementations, or paired regression tests. This release… See the full description on the dataset page: https://huggingface.co/datasets/skonml/code-contract-repair.texttext-generation1K<n<10K0 likes53 downloads2mo agoHugging Face17kkomyoeminaung /code-cpt-corpus Dataset Card for code-cpt-corpus Dataset Summary ဒီ dataset က code-cpt-corpus အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train data/train.jsonl Licensing Information ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/code-cpt-corpus.text100K<n<1M0 likes49 downloads3mo agoHugging Face18codecainecowboy /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/codecainecowboy/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M0 likes45 downloads5mo agoHugging Face19HlebYakhnitski /codecomplex CodeComplex Dataset Dataset Description CodeComplex consists of 4,200 Java codes submitted to programming competitions by human programmers and their complexity labels annotated by a group of algorithm experts. How to use it You can load and iterate through the dataset with the following two lines of code: from datasets import load_dataset ds = load_dataset("codeparrot/codecomplex", split="train") print(next(iter(ds))) Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/HlebYakhnitski/codecomplex.texttext-generation1K<n<10K0 likes39 downloads10mo agoHugging Face20code-critic-model /critic-sft-qwen-only critic-sft-qwen-only The Qwen-only SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Qwen-only and Qwen3-4B-Critic-SFT-Qwen-only, the Qwen-only arms of the corpus ablation in Table 3. Each record is one critique point: a Qwen3-Next-80B-A3B-Instruct trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it under the paper's high-level prompt. The main corpus, critic-sft-cwm-qwen, is… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-qwen-only.texttext-generation1K<n<10K0 likes36 downloads1mo agoHugging Face21code-critic-model /PRM_1541i Not the paper's DPO data. This is an earlier set of 1,541 preference pairs used for development DPO sweeps. The released critic Qwen3-4B-Critic-SFT-DPO from Steer, Don't Solve was trained on a different set of 1,409 pairs built with the same procedure. See the organization page for the artifacts behind the paper. PRM_1541i 1541 preference pairs for training a critic over coding-agent trajectories. Each example is a multi-turn agent transcript paired with two candidate… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/PRM_1541i.texttext-generation1K<n<10K0 likes25 downloads1mo agoHugging Face22codecompletedeployment /base_dataset_07_nov_23This dataset is 10% repo sampled dataset for selected languages. We applied a repo sample rate of 10%. e.g. if sample rate is 10% then we take 10% of all repos for a given language but include all files inside the repo. This was generated using our codecomplete/training/completions/datagen ./launch.sh \ --dataset-name bigcode/starcoderdata \ --subset c,cpp,go,java,javascript,typescript,python,ruby,scala,sql \ --sample-rate 0.01 \ --hf-token <HF_TOKEN> \ --output-dir… See the full description on the dataset page: https://huggingface.co/datasets/codecompletedeployment/base_dataset_07_nov_23.text100K<n<1M0 likes18 downloads3y agoHugging Face23codecomplete /base_datasetThis dataset is 10% repo sampled dataset for selected languages. We applied a repo sample rate of 10%. e.g. if sample rate is 10% then we take 10% of all repos for a given language but include all files inside the repo. This was generated using our codecomplete/training/completions/datagen ./launch.sh \ --dataset-name bigcode/starcoderdata \ --subset c,cpp,go,java,javascript,typescript,python,ruby,scala,sql \ --sample-rate 0.01 \ --hf-token <HF_TOKEN> \ --output-dir… See the full description on the dataset page: https://huggingface.co/datasets/codecomplete/base_dataset.text100K<n<1M0 likes17 downloads3y agoHugging Face24semeru /Code-Code-CloneDetection-POJ104 Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Clone-detection-POJ-104 in Semeru CodeXGLUE -- Clone Detection (POJ-104) Task Definition Given a code and a collection of candidates as the input, the task is to return Top K codes with the same semantic. Models are evaluated by MAP@R score. MAP@R is defined as the mean of… See the full description on the dataset page: https://huggingface.co/datasets/semeru/Code-Code-CloneDetection-POJ104.text10K<n<100K2 likes16 downloads4y agoHugging Face25semeru /Code-code-galeras-prompting-3k-treatment-1tabular1K<n<10K0 likes13 downloads3y agoHugging Face26devesh5 /codeconv-fortran-to-rusttexttranslationn<1K2 likes12 downloads3y agoHugging Face27macavaney /codectext100K<n<1M0 likes12 downloads2y agoHugging Face28semeru /Code-code-galeras-prompting-3k-treatment-2tabular1K<n<10K0 likes9 downloads3y agoHugging Face29anhnct /audio_codectext10K<n<100K0 likes9 downloads4mo agoHugging Face30Valliappan /CodeComp2text10K<n<100K1 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.