datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-smith-cppthe-stack-v2-new-cppSWE-Bench-MultilingualC_CPPFileteredSWE-Bench-MultilingualC_CPPFiletered_newstack-v2-cpp-2019the-stack-v2-cpparc-stack-cppLangMap-TheStack-cpp-100M
LangMap-TheStack-cpp-100M
Code finetuning dataset for cpp streamed from bigcode/the-stack.
Tokens collected: 100,000,000 (target: 100,000,000)
Tokenizer: allenai/OLMo-3-1025-7B
Schema: {"text": [...]} (sanitised source code)
code_contest_instruct_cppIDD_Detection_CPPE5The IDD Object Detection dataset containing 40K images with CPPE-5 like (or YOLO) dataset annotation format.Refer to the original dataset: https://idd.insaan.iiit.ac.in
codeforces_cpp_cleaned_scaled_classcpp_cwe_GRPO
cpp_cwe_GRPO
VeRL/GRPO-ready C++ security-coding dataset produced from every harness-passing stage-6 rewrite in the simple_gen pipeline.
Files
File
Rows
Description
cpp_cwe_GRPO.parquet
8,039
Full corrected dataset
cpp_cwe_GRPO_train.parquet
7,236
Deterministic 90% training split
cpp_cwe_GRPO_val.parquet
803
Deterministic 10% validation split
The split uses seed 20260929. Each record has a unique function name, and the two splits are disjoint.… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.stack_edu_cppSWE-smith-cppcpp_unit_tests_benchmark_datacpp-mit-github-search-code-in-reposStackless_CPP_V2leetcode-codegen-cpp
LeetCode Code-Gen Dataset — C++
4025 rows. Given a problem statement, its input/output examples, and a
required algorithm/technique, generate a correct C++ solution.
Part of a 4-language collection built from the same source: see the sibling
Python,
Java,
C++, and
JavaScript
datasets.
Verification
Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp.HPC_Fortran_CPPThis dataset is associated with the following paper:
Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++,
Links
https://arxiv.org/abs/2307.07686
https://github.com/bin123apple/OpenMP-Fortran-CPP-Translation
codeforces_cpp_cleanedthe-stack-v2-filtered-cppexp_rpt_nemotron-cpp-v2-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_nemotron-cpp-v2-qwen3.5-122b-131k-opencode-traces.gpt-5-mini-rebench-v2-cppnemotron-cpp-qwen3.5-122b-32k-tracesglm52-datagen-r11-36-nemotron-cpp-tracesproject_codeNet_translation_go_cpp
📂 Translation_go_into_cpp
Bienvenue sur la base de données Translation_go_into_cpp.
Ce dataset regroupe des traductions de code en trois langages de programmation : Go, Python et C++.Chaque ligne contient un script go, et une traduction de celui ci, soit en python, soit en C++.
L'objectif principal de ce dataset est de fournir une base propre et nettoyée pour l'entraînement
de modèles de type LLM (Large Language Models) dans des tâches de traduction Go ↔ C++.
📊… See the full description on the dataset page: https://huggingface.co/datasets/SprayOpoivre/project_codeNet_translation_go_cpp.rlvr-code-data-cppCPP-UNITTEST-BENCH
Dataset Card for Open Source Code and Unit Tests
Dataset Details
Dataset Description
This dataset contains c++ code snippets and their corresponding ground truth unit tests collected from various open-source GitHub repositories. The primary purpose of this dataset is to aid in the development and evaluation of automated testing tools, code quality analysis, and LLM models for test generation.
Curated by: Vaishnavi Bhargava
Language(s): C++… See the full description on the dataset page: https://huggingface.co/datasets/Nutanix/CPP-UNITTEST-BENCH.exp_rpt_nemotron-cpp-minimax-m27-131k-tracesCpp-Math
Cpp-Math: A Math-to-C++ Instruction-Tuning Dataset
Dataset Overview
The Cpp-Math dataset is designed for fine-tuning models to translate mathematical problems into C++ code. It focuses on evaluating the ability of language models to generate accurate and executable C++ code from mathematical expressions or problem statements. The dataset is particularly useful for benchmarking models on tasks that require both mathematical reasoning and programming skills.
Data… See the full description on the dataset page: https://huggingface.co/datasets/BoltzmannEntropy/Cpp-Math.
