Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CyberNative /Code_Vulnerability_Security_DPO Secure-code dataset: correction and audit notice (2026-10-06) The chosen and rejected fields are original model-generated preference labels. They are not verified secure/insecure classifications. Do not use these labels as security ground truth or use the dataset as a validated secure-code training or evaluation set without your own contextual validation. The historical description below overstates security, optimization, realism and validation. Those claims are superseded by… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.text1K<n<10K173 likes1.3k downloads4d agoHugging Face02CIRCL /vulnerability-cwe-patch Description This dataset, CIRCL/vulnerability-cwe-patch, provides structured, real-world vulnerabilities enriched with CWE identifiers and corresponding patches from platforms like GitHub and GitLab. It is designed to support the development of tools for vulnerability classification, triage, and automated remediation. Each entry includes metadata such as CVE/GHSA ID, a description, CWE categorization, and links to verified patch commits with associated diff content and commit… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-cwe-patch.text1K<n<10K5 likes517 downloads3mo agoHugging Face03ayshajavd /code-security-vulnerability-dataset Code Security Vulnerability Dataset A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories. Dataset Details Property Value Total Samples 175,419 Train / Val / Test 140,335 / 17,542 / 17,542 Languages C, C++, Python, JavaScript, Java, PHP, Go Labels 31 (multi-label) Format Parquet with… See the full description on the dataset page: https://huggingface.co/datasets/ayshajavd/code-security-vulnerability-dataset.texttext-classification100K<n<1M8 likes360 downloads6mo agoHugging Face04lemon42-ai /Code_Vulnerability_Labeled_Dataset Dataset Card for Code_Vulnerability_Labeled_Dataset Dataset Summary This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation: CWE Description CWE-020 Improper Input Validation CWE-022 Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”) CWE-078 Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”) CWE-079 Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.texttext-classification1K<n<10K13 likes302 downloads2y agoHugging Face05CIRCL /vulnerability-scores vulnerability-scores This dataset comprises 811,260 real-world vulnerabilities used to train and evaluate VLAI, a transformer-based model designed to predict software vulnerability severity levels directly from text descriptions, enabling faster and more consistent triage. The dataset is presented in the paper VLAI: A RoBERTa-Based Model for Automated Vulnerability Severity Classification. Sources Source Label Entries Share cvelistv5 CVE Program… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-scores.tabulartext-classification100K<n<1M11 likes296 downloads5d agoHugging Face06CIRCL /vulnerability-attack-techniques vulnerability-attack-techniques This dataset maps 1,207 CVEs to MITRE ATT&CK (Enterprise) techniques, joining hand-curated mappings from the MITRE Center for Threat-Informed Defense (CTID) with vulnerability descriptions from CIRCL/vulnerability-scores. It is intended for training and evaluating models that suggest candidate ATT&CK techniques from a vulnerability description: CVSS tells you how bad a vulnerability is, CWE what kind of flaw it is — ATT&CK tells defenders what… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques.texttext-classification1K<n<10K1 likes261 downloads2mo agoHugging Face07ryal-xyz /vul-mine-vulnerability-dataset VulMine vulnerability dataset VulMine is a strict, naturally imbalanced function/method-level vulnerability dataset mined from public OSV advisories and immutable public Git revisions. The default configuration covers Java, Python, and JavaScript. The c-cpp configuration is a C/C++ language control built with the same VulMine label and cleaning principles for comparisons with Big-Vul. Release The default configuration contains vulmine-clean-v1.2. Split… See the full description on the dataset page: https://huggingface.co/datasets/ryal-xyz/vul-mine-vulnerability-dataset.tabulartext-classification100K<n<1M1 likes244 downloads2mo agoHugging Face08cmonplz /Python_Vulnerability_Remediation Python SAST Vulnerability and Remediation Dataset Summary This dataset is a collection of Python code snippets containing common security vulnerabilities, paired with their corresponding high-quality remediations. It is designed for fine-tuning language models to assist with Static Analysis Security Testing (SAST) by suggesting secure code fixes. The dataset is primarily focused on vulnerabilities from the following Common Weakness Enumerations (CWEs): CWE-89 (SQL… See the full description on the dataset page: https://huggingface.co/datasets/cmonplz/Python_Vulnerability_Remediation.text1K<n<10K1 likes210 downloads11mo agoHugging Face09ismailtasdelen /unified-vulnerability-intelligence-dataset Unified Vulnerability Intelligence Dataset (UVID) v3.0 — Cyber Security Knowledge Graph UVID is a structured cyber security knowledge graph that unifies multiple vulnerability classification frameworks into a single knowledge base. Each of the 250 records describes one application/software security vulnerability and links it — where authoritative data exists — across CWE, CAPEC, MITRE ATT&CK, CVSS, 14 OWASP projects, secure-fix intelligence, detection surfaces, programming… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/unified-vulnerability-intelligence-dataset.texttext-classificationn<1K1 likes205 downloads3mo agoHugging Face10SecCoderX /SecCoderX_Reasoning_Vulnerability_Detection_SFT_Cold_Start_Dataset Citation If you find our work helpful, feel free to give us a cite. @misc{wu2026securecodegenerationonline, title={Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model}, author={Tianyi Wu and Mingzhe Du and Yue Liu and Chengran Yang and Terry Yue Zhuo and Jiaheng Zhang and See-Kiong Ng}, year={2026}, eprint={2602.07422}, archivePrefix={arXiv}, primaryClass={cs.CR}, url={https://arxiv.org/abs/2602.07422}… See the full description on the dataset page: https://huggingface.co/datasets/SecCoderX/SecCoderX_Reasoning_Vulnerability_Detection_SFT_Cold_Start_Dataset.text10K<n<100K0 likes173 downloads7mo agoHugging Face11pranay5255 /smart-contract-vulnerability-benchmarks Smart Contract Vulnerability Benchmarks smartbugs-wild/contracts is stored as lossless CSV shards with content_base64; other benchmark files are raw. Converted smartbugs-wild contracts: 47398 Raw files: 5295 Generated: 2026-06-24T16:06:36.674733+00:00 0 likes171 downloads4mo agoHugging Face12darkknight25 /Smart_Contract_Vulnerability_DatasetSmart Contract Vulnerability Dataset Overview The Smart Contract Vulnerability Dataset (SCV-1-2000) is a comprehensive JSONL dataset containing 2000 entries (SCV-1 to SCV-2000) focused on advanced and unconventional smart contract vulnerabilities and attack vectors, with an emphasis on Decentralized Finance (DeFi) protocols. This dataset is designed for cybersecurity professionals, blockchain developers, machine learning engineers, and data scientists to train models, evaluate… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Smart_Contract_Vulnerability_Dataset.text-classification1K<n<10K6 likes151 downloads1y agoHugging Face13msc-smart-contract-auditing /vulnerability-severity-classificationThis dataset combines vulnerable functions (scraped from 5 auditting companies: Codehawks, ConsenSys, Cyfrin, Sherlock, Trust Security) and auddited functions with no vulnerabilities (scraped from Etherscan) The purpose of the dataset is to enable training of classification models to discriminate between the 4 classes: none, low, medium and high. Field Description 1. function Raw solidity code 2. severity Severity of vulnerability ('none', low, medium, high) Data… See the full description on the dataset page: https://huggingface.co/datasets/msc-smart-contract-auditing/vulnerability-severity-classification.texttext-classification1K<n<10K3 likes150 downloads2y agoHugging Face14316usman /vulnerability-triage VULNERABILITY_TRIAGE A preference dataset for VULNERABILITY_TRIAGE, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/vulnerability-triage.texttext-generation1K<n<10K0 likes141 downloads12d agoHugging Face15giabaohuynhasu /cna-vulnerability-census-replication Empirical Vulnerability Census (1999–2026, $N = 385,524$), Cybernetic Queueing Instability, and CISA BOD 26-04 Remediation Deficit Deterministic Empirical Replication Package & Econometric Audits Principal Investigator: Gia Bao Huynh (Jun Huynh)ORCID: 0009-0008-2372-5852Affiliation: Independent Scholar / Ho Chi Minh City, VietnamLive Interactive Simulator: Cybernetic Queueing Instability Simulator (M/G/1) 🏛️ Executive Summary & Theoretical… See the full description on the dataset page: https://huggingface.co/datasets/giabaohuynhasu/cna-vulnerability-census-replication.texttabular-regression100K<n<1M0 likes130 downloads22d agoHugging Face16CIRCL /Vulnerability-FSTEC Vulnerability-FSTEC Vulnerability descriptions and severity labels from the FSTEC, extracted via Vulnerability-Lookup. Source Data source: Vulnerability-Lookup API Extraction tool: VulnTrain Related models CIRCL/vulnerability-severity-classification-russian-ruRoberta-large — severity classifier trained on this dataset text10K<n<100K0 likes128 downloads5mo agoHugging Face17CIRCL /vulnerability-attack-techniques-llm-scaling vulnerability-attack-techniques-llm-scaling ⚠️ The labels in this dataset are machine-generated by an LLM, not analyst-curated — and the paper that produced them found they do not improve a classifier trained on the expert gold set. It is published for reproducibility and for research on LLM-assisted labeling. For training, use the curated gold set CIRCL/vulnerability-attack-techniques. This dataset contains 984 CVEs labeled with MITRE ATT&CK (Enterprise) techniques by… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques-llm-scaling.texttext-classificationn<1K1 likes108 downloads2mo agoHugging Face18olukotunjosh /cve-vulnerability-ranking-results0 likes104 downloads3mo agoHugging Face19CIRCL /vulnerability Dataset Card for Dataset Name This dataset has been generated with: https://github.com/vulnerability-lookup/VulnTrain Based on data from the Vulnerability-Lookup instance operated by CIRCL: https://vulnerability.circl.lu/ The dataset is derived from CVE data provided by NIST and enriched with information from the CVE Program, FKIE, and Vulnrichment. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability.text100K<n<1M1 likes98 downloads1y agoHugging Face20davidquicast /vulnerability-intelligence-diagrammatic-reasoning Vulnerability Intelligence with Diagrammatic Reasoning [!Important] This dataset was created as a proof-of-concept for the Reasoning Datasets Competition (May 2025). If you have any feedback or suggestions, please feel free to open a discussion! Access the Github repository here. A. Overview This dataset focuses on security vulnerability analysis through a multi-dimensional approach that combines four types of reasoning to generate valuable insights for… See the full description on the dataset page: https://huggingface.co/datasets/davidquicast/vulnerability-intelligence-diagrammatic-reasoning.textn<1K1 likes85 downloads1y agoHugging Face21Leopo1d /OpenVul_Rejection_Sampling_based_Vulnerability_Reasoning_Dataset_for_SFTThis dataset provides high-quality, correctness-filtered vulnerability reasoning data to support the SFT of specialized VD LLMs for future research. text1K<n<10K1 likes85 downloads8mo agoHugging Face22ChamaraVishwajithRajapaksha /Code-Vulnerability-FineTune 🔐 Code Vulnerability FineTome — CWE-Enriched Conversation Dataset 📌 Overview This dataset converts raw security-labeled C/C++ code samples into instruction-following conversation pairs suitable for fine-tuning large language models (LLMs) on software vulnerability detection and analysis. It is built by preprocessing and transforming the ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset (330k rows, sourced from DiverseVul + MITRE CWE enrichment) into… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code-Vulnerability-FineTune.texttext-generation100K<n<1M0 likes84 downloads6mo agoHugging Face23CIRCL /Vulnerability-CNVD Vulnerability-CNVD Vulnerability descriptions and severity labels from the China National Vulnerability Database (CNVD), extracted via Vulnerability-Lookup. Dataset structure Field Type Description id string CNVD identifier (e.g., CNVD-2025-03529) title string Vulnerability title in Chinese description string Vulnerability description in Chinese severity string Severity level: 高 (High), 中 (Medium), or 低 (Low) cve_id string Corresponding CVE… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/Vulnerability-CNVD.text100K<n<1M4 likes83 downloads1mo agoHugging Face24Leopo1d /OpenVul_Distilled_Vulnerability_Reasoning_CoTs_from_DeepSeek-R1-0528This dataset provides all training data's vulnerability reasoning CoTs (with 8 generations per sample) distilled from DeepSeek-R1-0528. This dataset has not been filtered for correctness and can be used to construct vulnerability reasoning and preference datasets for future research. text10K<n<100K1 likes79 downloads8mo agoHugging Face25hzhu721 /vulnerability-specifications VulInstruct: Specification-Guided Vulnerability Detection 🔒 📄 Paper   |   🤗 Dataset   |   💻 GitHub Introduction This repository contains the Specification Knowledge Base introduced in the paper VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications. VulInstruct is a specification-guided approach that systematically extracts security specifications—expectations about how code should behave to remain safe—from… See the full description on the dataset page: https://huggingface.co/datasets/hzhu721/vulnerability-specifications.texttext-classification1K<n<10K1 likes66 downloads6mo agoHugging Face26leocyusa /ra-quant-vulnerability-maps Reasoning-Aware Quantization: Vulnerability Maps for DeepSeek-R1 Distilled Models Pre-computed, per-module INT4 vulnerability maps for DeepSeek-R1 distilled reasoning models, plus accuracy–energy results for selective mixed-precision compression. Use the maps to decide which layers to keep in higher precision without repeating the profiling (336 single-module evaluations per benchmark for a 14B model). Produced with compute provided by OpenToken, at Carnegie Mellon University… See the full description on the dataset page: https://huggingface.co/datasets/leocyusa/ra-quant-vulnerability-maps.0 likes60 downloads6d agoHugging Face27burpsuite /Code_Vulnerability_Security_DPO Cybernative.ai Code Vulnerability and Security Dataset Dataset Description The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/burpsuite/Code_Vulnerability_Security_DPO.text1K<n<10K2 likes58 downloads9mo agoHugging Face28AgileRLArena /vulnerability-scores-cvss-v3 Vulnerability scores (CVSS v3 combined) A labeled slice of CIRCL/vulnerability-scores for training and evaluating models that predict CVSS v3 severity from a vulnerability description. Every row has a combined v3 score and a severity band. Rows with no v3.1 or v3.0 score were dropped. What changed from the original The CIRCL dataset stores four separate CVSS columns (cvss_v4_0, cvss_v3_1, cvss_v3_0, cvss_v2_0). Those versions are not on the same scale, so this… See the full description on the dataset page: https://huggingface.co/datasets/AgileRLArena/vulnerability-scores-cvss-v3.tabulartext-classification100K<n<1M0 likes56 downloads1mo agoHugging Face29Leopo1d /OpenVul_Vulnerability_Query_Dataset_for_RLThis dataset provides context-aware vulnerability queries partitioned chronologically by commit date into training, validation, and test sets, designed to support the RL (e.g., GRPO) of specialized VD LLMs in future research. text10K<n<100K0 likes54 downloads8mo agoHugging Face30maddyrucos /code_vulnerability_pythontabularn<1K4 likes52 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.