datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols.
~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks.
The only human artifacts are:
agents/*
human/*
AGENTS.md
SWE-smith-cpphuman_eval_cppthe-stack-v2-new-cppSWE-Bench-MultilingualC_CPPFileteredSWE-Bench-MultilingualC_CPPFiletered_newCpp-Code-LargeCpp-Code-Large
Cpp-Code-Large is a large-scale corpus of C++ source code comprising more than 5 million lines of C++ code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the C++ ecosystem.
By providing a high-volume, language-specific corpus, Cpp-Code-Large enables systematic experimentation in C++-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Cpp-Code-Large.the-stack-v2-cppllama.cppversion https://git-lfs.github.com/spec/v1
oid sha256:cfc44b7ba25614df70e6b65e3341cae0310163bd32fd31a6b928a542df433faf
size 30786
stack-v2-cpp-2019LiveCodeBench-CPP
LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++
Overview
LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems).
AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.arc-stack-cppmagenta-realtime-mlx-cpp
Magenta RealTime — C++ MLX runtime bundle
This dataset is a re-packaging of
Google's Magenta RealTime weights
for the C++ MLX runtime in
rhymeswithlion/magenta-realtime-mlx-cpp.
It contains exactly what mlx-stream needs at startup; nothing more, nothing
less. The upstream .pt / .npy checkpoints are intentionally not
mirrored here — they're only useful for the (Python) re-export tooling on the
project's main distribution.
Contents
.
├──… See the full description on the dataset page: https://huggingface.co/datasets/rhymeswithlion/magenta-realtime-mlx-cpp.MD-trajectories-CPPF-tubulin-heterodimer-and-monomers
MD-trajectories-CPPF-tubulin-heterodimer-and-monomers
Copy this file into the Hugging Face dataset “README” (Dataset card).Source of truth in Git: https://github.com/jasperyeoh/integrative-ai-assisted-modeling-of-cppf-tubulin-interactions — see docs/DIMER_TRAJECTORY_NAMING.md.
What this dataset contains
All-atom GROMACS production trajectories (.xtc) for CPPF with human tubulin:
5IJ0 / soluble curved dimer (main text): three heterodimer replicates extended to… See the full description on the dataset page: https://huggingface.co/datasets/jasperyeoh2/MD-trajectories-CPPF-tubulin-heterodimer-and-monomers.SWE-smith-cppTopoBox-3D
TopoBox-3D
Paper (arXiv:2609.05860) | Code (GitHub)
TopoBox-3D is the dataset accompanying Beyond Arbitrary Geometry: Topology
Generalization in Neural PDE Operators. It is a controlled three-dimensional
benchmark for separating fixed-topology geometry shift from generalization to
unseen homological support.
The benchmark contains 5,280 connected box-minus-void geometries and 63,360
fixed-time Hodge-heat instances. Through-tunnels and enclosed cavities control
the first and… See the full description on the dataset page: https://huggingface.co/datasets/cppyyy/TopoBox-3D.code_contest_instruct_cppLangMap-TheStack-cpp-100M
LangMap-TheStack-cpp-100M
Code finetuning dataset for cpp streamed from bigcode/the-stack.
Tokens collected: 100,000,000 (target: 100,000,000)
Tokenizer: allenai/OLMo-3-1025-7B
Schema: {"text": [...]} (sanitised source code)
synthetic-cpp
Dataset Card for Synthetic C++ Dataset
Dataset Description
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [---
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [https://huggingface.co/datasets/ReySajju742/synthetic-cpp/]
Point of Contact: [ReySajju742]
Dataset Summary
This dataset contains 10,000 rows of synthetically generated data focusing on the topic of "C++… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/synthetic-cpp.IDD_Detection_CPPE5The IDD Object Detection dataset containing 40K images with CPPE-5 like (or YOLO) dataset annotation format.Refer to the original dataset: https://idd.insaan.iiit.ac.in
cpp_unit_tests_benchmark_datacpp-code-code_search_net-style
C++ Dataset
documentation source: https://huggingface.co/docs/datasets/main/en/repository_structure
Supported Tasks and Leaderboards
language-modeling: The dataset can be used to train a model for modelling programming languages, which consists in building language models for programming languages.
Language
C++ programming language
Dataset Structure
Data Instances
A data point consists of a function code along with its documentation.… See the full description on the dataset page: https://huggingface.co/datasets/malteklaes/cpp-code-code_search_net-style.stack_edu_cppcodeforces_cpp_cleaned_scaled_classthe-stack-v2-filtered-cppcpp_unit_tests_benchmark_data_with_splitscpp_cwe_GRPO
cpp_cwe_GRPO
VeRL/GRPO-ready C++ security-coding dataset produced from every harness-passing stage-6 rewrite in the simple_gen pipeline.
Files
File
Rows
Description
cpp_cwe_GRPO.parquet
8,039
Full corrected dataset
cpp_cwe_GRPO_train.parquet
7,236
Deterministic 90% training split
cpp_cwe_GRPO_val.parquet
803
Deterministic 10% validation split
The split uses seed 20260929. Each record has a unique function name, and the two splits are disjoint.… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.gpt-5-mini-rebench-v2-cppHPC_Fortran_CPPThis dataset is associated with the following paper:
Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++,
Links
https://arxiv.org/abs/2307.07686
https://github.com/bin123apple/OpenMP-Fortran-CPP-Translation
cpp-mit-github-search-code-in-repos
