datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code_Vulnerability_Security_DPO
Secure-code dataset: correction and audit notice (2026-10-06)
The chosen and rejected fields are original model-generated preference labels. They are not verified secure/insecure classifications. Do not use these labels as security ground truth or use the dataset as a validated secure-code training or evaluation set without your own contextual validation.
The historical description below overstates security, optimization, realism and validation. Those claims are superseded by… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.code-security-vulnerability-dataset
Code Security Vulnerability Dataset
A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories.
Dataset Details
Property
Value
Total Samples
175,419
Train / Val / Test
140,335 / 17,542 / 17,542
Languages
C, C++, Python, JavaScript, Java, PHP, Go
Labels
31 (multi-label)
Format
Parquet with… See the full description on the dataset page: https://huggingface.co/datasets/ayshajavd/code-security-vulnerability-dataset.Code_Vulnerability_Security_DPO
Cybernative.ai Code Vulnerability and Security Dataset
Dataset Description
The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/burpsuite/Code_Vulnerability_Security_DPO.Python-Security-Code-DatasetCWE-Code_Vulnerability_Security_DPOCyberNative_Code_Vulnerability_Security_DPO-PreferenceShareGPTai-code-security-golden
AI Code Security — Golden Set
A small, hand-labeled benchmark of code snippets — vulnerable, safe, and
needs-audit — for evaluating how well a tool detects security problems in
AI-generated code ("vibe coding"). Every case is a minimal, self-contained
example with a known, by-construction ground-truth label.
Crucially, the set is built around safe twins: many vulnerable cases are
paired with a near-identical safe variant living at the same file path. This
makes the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/axyr/ai-code-security-golden.Code_Vulnerability_Security_DPOsecurity-code-chatbot_pre-trainCode_Vulnerability_Security_DPO-archive
Cybernative.ai Code Vulnerability and Security Dataset
Dataset Description
The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Code_Vulnerability_Security_DPO-archive.mirror-code-security-vulnerability-dataset
Code Security Vulnerability Dataset
A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories.
Dataset Details
Property
Value
Total Samples
175,419
Train / Val / Test
140,335 / 17,542 / 17,542
Languages
C, C++, Python, JavaScript, Java, PHP, Go
Labels
31 (multi-label)
Format
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-code-security-vulnerability-dataset.mirror-Code_Vulnerability_Security_DPO
Cybernative.ai Code Vulnerability and Security Dataset
Dataset Description
The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Code_Vulnerability_Security_DPO.cybernative-code-vulnerability-security-dpo
Mirror: CyberNative/Code_Vulnerability_Security_DPO
Pinned snapshot / mirror of CyberNative/Code_Vulnerability_Security_DPO, re-hosted for PROTISEC
research reproducibility. Redistributed under the upstream license (apache-2.0)
with attribution — all credit to the original author.
Original author: CyberNative
Source dataset: CyberNative/Code_Vulnerability_Security_DPO
License: apache-2.0
Family: coding_security
Mode: full
Rows cached: 4657
Changes vs upstream: cached… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/cybernative-code-vulnerability-security-dpo.code-security-datasetcode-security
