Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SWE-bench /SWE-smith-cpptext1K<n<10K0 likes2.6k downloads7mo agoHugging Face02Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes766 downloads2y agoHugging Face03Romoamigo /SWE-Bench-MultilingualC_CPPFileteredtextn<1K0 likes522 downloads1y agoHugging Face04Romoamigo /SWE-Bench-MultilingualC_CPPFiletered_newtextn<1K0 likes487 downloads1y agoHugging Face05ningani /stack-v2-cpp-2019tabular10M<n<100M0 likes458 downloads2y agoHugging Face06Reset23 /the-stack-v2-cpptabular1M<n<10M1 likes400 downloads2y agoHugging Face07ThomasTheMaker /arc-stack-cpptabular1M<n<10M0 likes332 downloads11mo agoHugging Face08MultilingualUnigramLM /LangMap-TheStack-cpp-100M LangMap-TheStack-cpp-100M Code finetuning dataset for cpp streamed from bigcode/the-stack. Tokens collected: 100,000,000 (target: 100,000,000) Tokenizer: allenai/OLMo-3-1025-7B Schema: {"text": [...]} (sanitised source code) text10K<n<100K0 likes223 downloads6mo agoHugging Face09wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes214 downloads2y agoHugging Face10izzako /IDD_Detection_CPPE5The IDD Object Detection dataset containing 40K images with CPPE-5 like (or YOLO) dataset annotation format.Refer to the original dataset: https://idd.insaan.iiit.ac.in imageobject-detection10K<n<100K1 likes212 downloads1y agoHugging Face11nguyentruong-ins /codeforces_cpp_cleaned_scaled_classtext1M<n<10M0 likes164 downloads3y agoHugging Face12AetherPrior /cpp_cwe_GRPO cpp_cwe_GRPO VeRL/GRPO-ready C++ security-coding dataset produced from every harness-passing stage-6 rewrite in the simple_gen pipeline. Files File Rows Description cpp_cwe_GRPO.parquet 8,039 Full corrected dataset cpp_cwe_GRPO_train.parquet 7,236 Deterministic 90% training split cpp_cwe_GRPO_val.parquet 803 Deterministic 10% validation split The split uses seed 20260929. Each record has a unique function name, and the two splits are disjoint.… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.texttext-generation1K<n<10K0 likes132 downloads11d agoHugging Face13hongliu9903 /stack_edu_cpptabular10M<n<100M0 likes129 downloads1y agoHugging Face14shanxianzheng /SWE-smith-cpptext1K<n<10K0 likes121 downloads2mo agoHugging Face15CPP-UT-BENCH /cpp_unit_tests_benchmark_datatext1K<n<10K5 likes120 downloads2y agoHugging Face16Wholesomeisland /cpp-mit-github-search-code-in-repostext100K<n<1M0 likes120 downloads11mo agoHugging Face17AISE-TUDelft /Stackless_CPP_V2tabular100K<n<1M0 likes108 downloads1y agoHugging Face18AmareshHebbar /leetcode-codegen-cpp LeetCode Code-Gen Dataset — C++ 4025 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct C++ solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp.texttext-generation1K<n<10K1 likes108 downloads3mo agoHugging Face19HPC-Forran2Cpp /HPC_Fortran_CPPThis dataset is associated with the following paper: Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++, Links https://arxiv.org/abs/2307.07686 https://github.com/bin123apple/OpenMP-Fortran-CPP-Translation textn<1K8 likes104 downloads2y agoHugging Face20nguyentruong-ins /codeforces_cpp_cleanedtext1M<n<10M0 likes99 downloads3y agoHugging Face21Reset23 /the-stack-v2-filtered-cpptabular100K<n<1M0 likes90 downloads1y agoHugging Face22open-athena /exp_rpt_nemotron-cpp-v2-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_nemotron-cpp-v2-qwen3.5-122b-131k-opencode-traces.textn<1K0 likes76 downloads3mo agoHugging Face23beranki /gpt-5-mini-rebench-v2-cpptabularn<1K0 likes74 downloads5mo agoHugging Face24open-athena /nemotron-cpp-qwen3.5-122b-32k-tracestext1K<n<10K0 likes71 downloads3mo agoHugging Face25open-athena /glm52-datagen-r11-36-nemotron-cpp-tracestextn<1K0 likes68 downloads3mo agoHugging Face26SprayOpoivre /project_codeNet_translation_go_cpp 📂 Translation_go_into_cpp Bienvenue sur la base de données Translation_go_into_cpp. Ce dataset regroupe des traductions de code en trois langages de programmation : Go, Python et C++.Chaque ligne contient un script go, et une traduction de celui ci, soit en python, soit en C++. L'objectif principal de ce dataset est de fournir une base propre et nettoyée pour l'entraînement de modèles de type LLM (Large Language Models) dans des tâches de traduction Go ↔ C++. 📊… See the full description on the dataset page: https://huggingface.co/datasets/SprayOpoivre/project_codeNet_translation_go_cpp.text1M<n<10M0 likes67 downloads7mo agoHugging Face27saurabh5 /rlvr-code-data-cpptext100K<n<1M0 likes62 downloads1y agoHugging Face28Nutanix /CPP-UNITTEST-BENCH Dataset Card for Open Source Code and Unit Tests Dataset Details Dataset Description This dataset contains c++ code snippets and their corresponding ground truth unit tests collected from various open-source GitHub repositories. The primary purpose of this dataset is to aid in the development and evaluation of automated testing tools, code quality analysis, and LLM models for test generation. Curated by: Vaishnavi Bhargava Language(s): C++… See the full description on the dataset page: https://huggingface.co/datasets/Nutanix/CPP-UNITTEST-BENCH.text1K<n<10K3 likes59 downloads2y agoHugging Face29open-athena /exp_rpt_nemotron-cpp-minimax-m27-131k-tracestext1K<n<10K0 likes55 downloads4mo agoHugging Face30BoltzmannEntropy /Cpp-Math Cpp-Math: A Math-to-C++ Instruction-Tuning Dataset Dataset Overview The Cpp-Math dataset is designed for fine-tuning models to translate mathematical problems into C++ code. It focuses on evaluating the ability of language models to generate accurate and executable C++ code from mathematical expressions or problem statements. The dataset is particularly useful for benchmarking models on tasks that require both mathematical reasoning and programming skills. Data… See the full description on the dataset page: https://huggingface.co/datasets/BoltzmannEntropy/Cpp-Math.text1K<n<10K2 likes54 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.