vulnerability
vulnerability-severity-classification-roberta-basevulnerability-attack-technique-classification-roberta-basevulnerability-attack-technique-biencodervulnerability-severity-classification-chinese-macbert-basevulnerability-severity-classification-russian-ruRoberta-largecodebert-devign-code-vulnerability-detectorvulnerabilityDetection-StarEncoder-DevignSecCoderX_Reasoning_Vulnerability_Detection_Reward_Model
Code_Vulnerability_Security_DPO
Secure-code dataset: correction and audit notice (2026-10-06)
The chosen and rejected fields are original model-generated preference labels. They are not verified secure/insecure classifications. Do not use these labels as security ground truth or use the dataset as a validated secure-code training or evaluation set without your own contextual validation.
The historical description below overstates security, optimization, realism and validation. Those claims are superseded by… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.vulnerability-cwe-patch
Description
This dataset, CIRCL/vulnerability-cwe-patch, provides structured, real-world vulnerabilities enriched with CWE identifiers and corresponding patches from platforms like GitHub and GitLab. It is designed to support the development of tools for vulnerability classification, triage, and automated remediation. Each entry includes metadata such as CVE/GHSA ID, a description, CWE categorization, and links to verified patch commits with associated diff content and commit… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-cwe-patch.code-security-vulnerability-dataset
Code Security Vulnerability Dataset
A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories.
Dataset Details
Property
Value
Total Samples
175,419
Train / Val / Test
140,335 / 17,542 / 17,542
Languages
C, C++, Python, JavaScript, Java, PHP, Go
Labels
31 (multi-label)
Format
Parquet with… See the full description on the dataset page: https://huggingface.co/datasets/ayshajavd/code-security-vulnerability-dataset.Code_Vulnerability_Labeled_Dataset
Dataset Card for Code_Vulnerability_Labeled_Dataset
Dataset Summary
This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation:
CWE
Description
CWE-020
Improper Input Validation
CWE-022
Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”)
CWE-078
Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”)
CWE-079
Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.vulnerability-scores
vulnerability-scores
This dataset comprises 811,260 real-world vulnerabilities used to train and evaluate VLAI,
a transformer-based model designed to predict software vulnerability severity levels directly from text descriptions,
enabling faster and more consistent triage.
The dataset is presented in the paper VLAI: A RoBERTa-Based Model for Automated Vulnerability Severity Classification.
Sources
Source
Label
Entries
Share
cvelistv5
CVE Program… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-scores.vulnerability-attack-techniques
vulnerability-attack-techniques
This dataset maps 1,207 CVEs to MITRE ATT&CK (Enterprise) techniques, joining
hand-curated mappings from the MITRE Center for Threat-Informed Defense (CTID)
with vulnerability descriptions from
CIRCL/vulnerability-scores.
It is intended for training and evaluating models that suggest candidate ATT&CK
techniques from a vulnerability description: CVSS tells you how bad a
vulnerability is, CWE what kind of flaw it is — ATT&CK tells defenders what… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques.
