DanCip/github-issues-vul-detection-post
π‘οΈ GitHub Issues Vulnerability Detection A benchmark dataset for automating the detection of code vulnerabilities by analyzing GitHub Issues. Curated to validate the findings in Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues (2025). π Dataset Description GitHub Issues Vulnerability Detection is a specialized dataset designed to evaluate the feasibility of identifying software vulnerabilities early by analyzing textual discussions in GitHubβ¦ See the full description on the dataset page: https://huggingface.co/datasets/DanCip/github-issues-vul-detection-post.
<h1 align="center">π‘οΈ GitHub Issues Vulnerability Detection</h1>
<div align="center">
  
A benchmark dataset for automating the detection of code vulnerabilities by analyzing GitHub Issues. Curated to validate the findings in [Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues](https://arxiv.org/abs/2501.05258) (2025).
</div>
π Dataset Description
GitHub Issues Vulnerability Detection is a specialized dataset designed to evaluate the feasibility of identifying software vulnerabilities early by analyzing textual discussions in GitHub issues.
Traditional vulnerability detection relies heavily on static code analysis or community reporting after patches are applied. However, early indicators of vulnerabilities often appear in informal communication channels like GitHub issues before they are officially recognized. This dataset targets this "pre-disclosure" window by linking real-world GitHub discussions to confirmed CVEs.
This dataset was created to support research into Transformer-based vulnerability detection, enabling models to distinguish between standard bug reports and critical security flaws based solely on textual descriptions.
β‘ Key Features
- Novelty: The first dataset to map GitHub issues directly to CVE records for automated classification.
- Real-World Data: Sourced from 31 top open-source repositories known for high vulnerability tracking activity.
- Balanced Context: Includes both vulnerability-related issues (positive samples) and standard non-security bugs (negative samples) to reflect realistic noise ratios.
- Rich Metadata: Contains full CVE metrics (CVSS scores), issue descriptions, and embeddings.
- Time-Aware Split: Specifically split to avoid data leakage regarding model training cutoffs (post-September 2021).
π Dataset Structure
Data Instances
Each instance represents a GitHub issue paired with ground truth labels indicating whether it is linked to a confirmed CVE vulnerability.
π Data Fields
This dataset includes comprehensive metadata from both the National Vulnerability Database (NVD) and GitHub API.
π οΈ Dataset Creation
Curation Rationale
Zero-day vulnerabilities and embargoed flaws are often discussed as "bugs" before they are officially labeled as vulnerabilities. Detecting these early can significantly reduce the window of exploitation. This dataset enables the training of LLMs and classifiers to flag these high-risk issues automatically.
Source Data
- Primary Source: GitHub Issues from open-source repositories.
- Ground Truth: National Vulnerability Database (NVD).
- Selection Criteria:
- Repositories: Ranked by the number of associated vulnerabilities; the top 31 were selected.
- Linkage: Issues were linked to CVEs via external references in NVD entries.
- Length: Issues exceeding 8,191 tokens were excluded to fit standard LLM context windows.
Dataset Splits
To ensure fair evaluation against models like GPT-3.5-Turbo (which has a training cutoff of Sept 2021), the dataset is split chronologically:
- Post-Cutoff Train: Issues created after the cutoff, used for training classifiers and fine-tuning.
- Post-Cutoff Test: Issues created after the cutoff, used exclusively for evaluation.
- Note: Issues pre-dating September 2021 were excluded to prevent data contamination.
π Citation
If you use this dataset, please cite the original paper:
@article{cipollone2025automating,
title={Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues},
author={Cipollone, Daniele and Izadi, Maliheh and Wang, Changjie and Scazzariello, Mariano and Ferlin, Simone and KostiΔ, Dejan and Chiesa, Marco},
journal={arXiv preprint arXiv:2501.05258},
year={2025}
}
