Team Ai
Datasetpublic

DanCip/github-issues-vul-detection-post

πŸ›‘οΈ GitHub Issues Vulnerability Detection A benchmark dataset for automating the detection of code vulnerabilities by analyzing GitHub Issues. Curated to validate the findings in Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues (2025). πŸ“– Dataset Description GitHub Issues Vulnerability Detection is a specialized dataset designed to evaluate the feasibility of identifying software vulnerabilities early by analyzing textual discussions in GitHub… See the full description on the dataset page: https://huggingface.co/datasets/DanCip/github-issues-vul-detection-post.

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes52downloads
Dataset Card

<h1 align="center">πŸ›‘οΈ GitHub Issues Vulnerability Detection</h1>

<div align="center">

![Paper](https://arxiv.org/abs/2501.05258) ![License](https://creativecommons.org/licenses/by/4.0/) ![Python](https://www.python.org/)

A benchmark dataset for automating the detection of code vulnerabilities by analyzing GitHub Issues. Curated to validate the findings in [Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues](https://arxiv.org/abs/2501.05258) (2025).

</div>


πŸ“– Dataset Description

GitHub Issues Vulnerability Detection is a specialized dataset designed to evaluate the feasibility of identifying software vulnerabilities early by analyzing textual discussions in GitHub issues.

Traditional vulnerability detection relies heavily on static code analysis or community reporting after patches are applied. However, early indicators of vulnerabilities often appear in informal communication channels like GitHub issues before they are officially recognized. This dataset targets this "pre-disclosure" window by linking real-world GitHub discussions to confirmed CVEs.

This dataset was created to support research into Transformer-based vulnerability detection, enabling models to distinguish between standard bug reports and critical security flaws based solely on textual descriptions.

⚑ Key Features

  • β€”Novelty: The first dataset to map GitHub issues directly to CVE records for automated classification.
  • β€”Real-World Data: Sourced from 31 top open-source repositories known for high vulnerability tracking activity.
  • β€”Balanced Context: Includes both vulnerability-related issues (positive samples) and standard non-security bugs (negative samples) to reflect realistic noise ratios.
  • β€”Rich Metadata: Contains full CVE metrics (CVSS scores), issue descriptions, and embeddings.
  • β€”Time-Aware Split: Specifically split to avoid data leakage regarding model training cutoffs (post-September 2021).

πŸ“‚ Dataset Structure

Data Instances

Each instance represents a GitHub issue paired with ground truth labels indicating whether it is linked to a confirmed CVE vulnerability.

πŸ“Š Data Fields

This dataset includes comprehensive metadata from both the National Vulnerability Database (NVD) and GitHub API.

FieldTypeDescription
issue_github_idint64Unique GitHub identifier for the issue.
issue_titlestringThe title of the GitHub issue.
issue_bodystringThe full textual content/description of the issue.
issue_msgstringThe processed message used for model input (title + body).
labelboolTrue if the issue is a vulnerability; False otherwise.
cve_idstringThe CVE identifier (e.g., CVE-2023-XXXX) if applicable.
cve_descriptionsstringOfficial description of the vulnerability from the NVD.
cve_publishedtimestampDate the CVE was officially published.
cve_metricsstructDetailed CVSS scoring metrics (V2, V3.1) assessing severity.
issue_created_attimestampDate the GitHub issue was created.
issue_owner_reposequenceThe [owner, repo] pair for the source repository.
issue_embeddingsequencePre-computed embeddings for the issue text.
issue_msg_n_tokensint64Token count of the issue message.

πŸ› οΈ Dataset Creation

Curation Rationale

Zero-day vulnerabilities and embargoed flaws are often discussed as "bugs" before they are officially labeled as vulnerabilities. Detecting these early can significantly reduce the window of exploitation. This dataset enables the training of LLMs and classifiers to flag these high-risk issues automatically.

Source Data

  • β€”Primary Source: GitHub Issues from open-source repositories.
  • β€”Ground Truth: National Vulnerability Database (NVD).
  • β€”Selection Criteria:
  • β€”Repositories: Ranked by the number of associated vulnerabilities; the top 31 were selected.
  • β€”Linkage: Issues were linked to CVEs via external references in NVD entries.
  • β€”Length: Issues exceeding 8,191 tokens were excluded to fit standard LLM context windows.

Dataset Splits

To ensure fair evaluation against models like GPT-3.5-Turbo (which has a training cutoff of Sept 2021), the dataset is split chronologically:

  • β€”Post-Cutoff Train: Issues created after the cutoff, used for training classifiers and fine-tuning.
  • β€”Post-Cutoff Test: Issues created after the cutoff, used exclusively for evaluation.
  • β€”Note: Issues pre-dating September 2021 were excluded to prevent data contamination.

πŸ“š Citation

If you use this dataset, please cite the original paper:

bibtex
@article{cipollone2025automating,
  title={Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues},
  author={Cipollone, Daniele and Izadi, Maliheh and Wang, Changjie and Scazzariello, Mariano and Ferlin, Simone and Kostić, Dejan and Chiesa, Marco},
  journal={arXiv preprint arXiv:2501.05258},
  year={2025}
}