datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cppe-5
Dataset Card for CPPE - 5
Dataset Summary
CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories.
Some features of this dataset are:
high quality images and annotations (~4.6 bounding boxes per image)
real-life images unlike any current such dataset
majority… See the full description on the dataset page: https://huggingface.co/datasets/rishitdagli/cppe-5.SWE-smith-cppthe-stack-v2-new-cppSWE-Bench-MultilingualC_CPPFileteredSWE-Bench-MultilingualC_CPPFiletered_newstack-v2-cpp-2019the-stack-v2-cpparc-stack-cppLangMap-TheStack-cpp-100M
LangMap-TheStack-cpp-100M
Code finetuning dataset for cpp streamed from bigcode/the-stack.
Tokens collected: 100,000,000 (target: 100,000,000)
Tokenizer: allenai/OLMo-3-1025-7B
Schema: {"text": [...]} (sanitised source code)
code_contest_instruct_cppIDD_Detection_CPPE5The IDD Object Detection dataset containing 40K images with CPPE-5 like (or YOLO) dataset annotation format.Refer to the original dataset: https://idd.insaan.iiit.ac.in
cppe-5-samplecodeforces_cpp_cleaned_scaled_classcpp_cwe_GRPO
cpp_cwe_GRPO
VeRL/GRPO-ready C++ security-coding dataset produced from every harness-passing stage-6 rewrite in the simple_gen pipeline.
Files
File
Rows
Description
cpp_cwe_GRPO.parquet
8,039
Full corrected dataset
cpp_cwe_GRPO_train.parquet
7,236
Deterministic 90% training split
cpp_cwe_GRPO_val.parquet
803
Deterministic 10% validation split
The split uses seed 20260929. Each record has a unique function name, and the two splits are disjoint.… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.stack_edu_cppSWE-smith-cppcpp_unit_tests_benchmark_datacpp-mit-github-search-code-in-reposStackless_CPP_V2leetcode-codegen-cpp
LeetCode Code-Gen Dataset — C++
4025 rows. Given a problem statement, its input/output examples, and a
required algorithm/technique, generate a correct C++ solution.
Part of a 4-language collection built from the same source: see the sibling
Python,
Java,
C++, and
JavaScript
datasets.
Verification
Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp.HPC_Fortran_CPPThis dataset is associated with the following paper:
Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++,
Links
https://arxiv.org/abs/2307.07686
https://github.com/bin123apple/OpenMP-Fortran-CPP-Translation
codeforces_cpp_cleanedthe-stack-v2-filtered-cppexp_rpt_nemotron-cpp-v2-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_nemotron-cpp-v2-qwen3.5-122b-131k-opencode-traces.gpt-5-mini-rebench-v2-cppnemotron-cpp-qwen3.5-122b-32k-tracesglm52-datagen-r11-36-nemotron-cpp-tracesproject_codeNet_translation_go_cpp
📂 Translation_go_into_cpp
Bienvenue sur la base de données Translation_go_into_cpp.
Ce dataset regroupe des traductions de code en trois langages de programmation : Go, Python et C++.Chaque ligne contient un script go, et une traduction de celui ci, soit en python, soit en C++.
L'objectif principal de ce dataset est de fournir une base propre et nettoyée pour l'entraînement
de modèles de type LLM (Large Language Models) dans des tâches de traduction Go ↔ C++.
📊… See the full description on the dataset page: https://huggingface.co/datasets/SprayOpoivre/project_codeNet_translation_go_cpp.rlvr-code-data-cppCPP-UNITTEST-BENCH
Dataset Card for Open Source Code and Unit Tests
Dataset Details
Dataset Description
This dataset contains c++ code snippets and their corresponding ground truth unit tests collected from various open-source GitHub repositories. The primary purpose of this dataset is to aid in the development and evaluation of automated testing tools, code quality analysis, and LLM models for test generation.
Curated by: Vaishnavi Bhargava
Language(s): C++… See the full description on the dataset page: https://huggingface.co/datasets/Nutanix/CPP-UNITTEST-BENCH.
