DanCip/github-issues-vul-detection-post
๐ก๏ธ GitHub Issues Vulnerability Detection A benchmark dataset for automating the detection of code vulnerabilities by analyzing GitHub Issues. Curated to validate the findings in Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues (2025). ๐ Dataset Description GitHub Issues Vulnerability Detection is a specialized dataset designed to evaluate the feasibility of identifying software vulnerabilities early by analyzing textual discussions in GitHubโฆ See the full description on the dataset page: https://huggingface.co/datasets/DanCip/github-issues-vul-detection-post.
051
1---2dataset_info:3 features:4 - name: cve_id5 dtype: string6 - name: cve_published7 dtype: timestamp[ns]8 - name: cve_descriptions9 dtype: string10 - name: cve_metrics11 struct:12 - name: cvssMetricV213 list:14 - name: acInsufInfo15 dtype: bool16 - name: baseSeverity17 dtype: string18 - name: cvssData19 struct:20 - name: accessComplexity21 dtype: string22 - name: accessVector23 dtype: string24 - name: authentication25 dtype: string26 - name: availabilityImpact27 dtype: string28 - name: baseScore29 dtype: float6430 - name: confidentialityImpact31 dtype: string32 - name: integrityImpact33 dtype: string34 - name: vectorString35 dtype: string36 - name: version37 dtype: string38 - name: exploitabilityScore39 dtype: float6440 - name: impactScore41 dtype: float6442 - name: obtainAllPrivilege43 dtype: bool44 - name: obtainOtherPrivilege45 dtype: bool46 - name: obtainUserPrivilege47 dtype: bool48 - name: source49 dtype: string50 - name: type51 dtype: string52 - name: userInteractionRequired53 dtype: bool54 - name: cvssMetricV3055 dtype: 'null'56 - name: cvssMetricV3157 list:58 - name: cvssData59 struct:60 - name: attackComplexity61 dtype: string62 - name: attackVector63 dtype: string64 - name: availabilityImpact65 dtype: string66 - name: baseScore67 dtype: float6468 - name: baseSeverity69 dtype: string70 - name: confidentialityImpact71 dtype: string72 - name: integrityImpact73 dtype: string74 - name: privilegesRequired75 dtype: string76 - name: scope77 dtype: string78 - name: userInteraction79 dtype: string80 - name: vectorString81 dtype: string82 - name: version83 dtype: string84 - name: exploitabilityScore85 dtype: float6486 - name: impactScore87 dtype: float6488 - name: source89 dtype: string90 - name: type91 dtype: string92 - name: cve_references93 list:94 - name: source95 dtype: string96 - name: tags97 sequence: string98 - name: url99 dtype: string100 - name: cve_configurations101 list:102 - name: nodes103 list:104 - name: cpeMatch105 list:106 - name: criteria107 dtype: string108 - name: matchCriteriaId109 dtype: string110 - name: versionEndExcluding111 dtype: string112 - name: versionEndIncluding113 dtype: string114 - name: versionStartExcluding115 dtype: 'null'116 - name: versionStartIncluding117 dtype: string118 - name: vulnerable119 dtype: bool120 - name: negate121 dtype: bool122 - name: operator123 dtype: string124 - name: operator125 dtype: string126 - name: url127 dtype: string128 - name: cve_tags129 sequence: string130 - name: domain131 dtype: string132 - name: issue_owner_repo133 sequence: string134 - name: issue_body135 dtype: string136 - name: issue_title137 dtype: string138 - name: issue_comments_url139 dtype: string140 - name: issue_comments_count141 dtype: int64142 - name: issue_created_at143 dtype: timestamp[ns]144 - name: issue_updated_at145 dtype: string146 - name: issue_html_url147 dtype: string148 - name: issue_github_id149 dtype: int64150 - name: issue_number151 dtype: int64152 - name: label153 dtype: bool154 - name: issue_msg155 dtype: string156 - name: issue_msg_n_tokens157 dtype: int64158 - name: issue_embedding159 sequence: float64160 splits:161 - name: post_train162 num_bytes: 62429952163 num_examples: 2166164 - name: post_test165 num_bytes: 41396204166 num_examples: 1445167 download_size: 71766293168 dataset_size: 103826156169configs:170- config_name: default171 data_files:172 - split: post_train173 path: data/post_train-*174 - split: post_test175 path: data/post_test-*176---177 178<h1 align="center">๐ก๏ธ GitHub Issues Vulnerability Detection</h1>179 180<div align="center">181 182[](https://arxiv.org/abs/2501.05258)183[](https://creativecommons.org/licenses/by/4.0/)184[](https://www.python.org/)185 186**A benchmark dataset for automating the detection of code vulnerabilities by analyzing GitHub Issues.**187*Curated to validate the findings in [Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues](https://arxiv.org/abs/2501.05258) (2025).*188 189</div>190 191---192 193## ๐ Dataset Description194 195**GitHub Issues Vulnerability Detection** is a specialized dataset designed to evaluate the feasibility of identifying software vulnerabilities early by analyzing textual discussions in GitHub issues.196 197Traditional vulnerability detection relies heavily on static code analysis or community reporting after patches are applied. However, early indicators of vulnerabilities often appear in informal communication channels like GitHub issues before they are officially recognized. This dataset targets this "pre-disclosure" window by linking real-world GitHub discussions to confirmed CVEs.198 199This dataset was created to support research into **Transformer-based vulnerability detection**, enabling models to distinguish between standard bug reports and critical security flaws based solely on textual descriptions.200 201### โก Key Features202* **Novelty:** The first dataset to map GitHub issues directly to CVE records for automated classification.203* **Real-World Data:** Sourced from 31 top open-source repositories known for high vulnerability tracking activity.204* **Balanced Context:** Includes both vulnerability-related issues (positive samples) and standard non-security bugs (negative samples) to reflect realistic noise ratios.205* **Rich Metadata:** Contains full CVE metrics (CVSS scores), issue descriptions, and embeddings.206* **Time-Aware Split:** Specifically split to avoid data leakage regarding model training cutoffs (post-September 2021).207 208---209 210## ๐ Dataset Structure211 212### Data Instances213Each instance represents a **GitHub issue** paired with ground truth labels indicating whether it is linked to a confirmed CVE vulnerability.214 215### ๐ Data Fields216 217This dataset includes comprehensive metadata from both the National Vulnerability Database (NVD) and GitHub API.218 219| Field | Type | Description |220| :--- | :--- | :--- |221| `issue_github_id` | `int64` | Unique GitHub identifier for the issue. |222| `issue_title` | `string` | The title of the GitHub issue. |223| `issue_body` | `string` | The full textual content/description of the issue. |224| `issue_msg` | `string` | The processed message used for model input (title + body). |225| `label` | `bool` | **True** if the issue is a vulnerability; **False** otherwise. |226| `cve_id` | `string` | The CVE identifier (e.g., CVE-2023-XXXX) if applicable. |227| `cve_descriptions` | `string` | Official description of the vulnerability from the NVD. |228| `cve_published` | `timestamp` | Date the CVE was officially published. |229| `cve_metrics` | `struct` | Detailed CVSS scoring metrics (V2, V3.1) assessing severity. |230| `issue_created_at` | `timestamp` | Date the GitHub issue was created. |231| `issue_owner_repo` | `sequence` | The `[owner, repo]` pair for the source repository. |232| `issue_embedding` | `sequence` | Pre-computed embeddings for the issue text. |233| `issue_msg_n_tokens`| `int64` | Token count of the issue message. |234 235---236 237## ๐ ๏ธ Dataset Creation238 239### Curation Rationale240Zero-day vulnerabilities and embargoed flaws are often discussed as "bugs" before they are officially labeled as vulnerabilities. Detecting these early can significantly reduce the window of exploitation. This dataset enables the training of LLMs and classifiers to flag these high-risk issues automatically.241 242### Source Data243* **Primary Source:** GitHub Issues from open-source repositories.244* **Ground Truth:** [National Vulnerability Database (NVD)](https://nvd.nist.gov/).245* **Selection Criteria:**246 * **Repositories:** Ranked by the number of associated vulnerabilities; the top 31 were selected.247 * **Linkage:** Issues were linked to CVEs via external references in NVD entries.248 * **Length:** Issues exceeding 8,191 tokens were excluded to fit standard LLM context windows.249 250### Dataset Splits251To ensure fair evaluation against models like GPT-3.5-Turbo (which has a training cutoff of Sept 2021), the dataset is split chronologically:252* **Post-Cutoff Train:** Issues created after the cutoff, used for training classifiers and fine-tuning.253* **Post-Cutoff Test:** Issues created after the cutoff, used exclusively for evaluation.254* *Note: Issues pre-dating September 2021 were excluded to prevent data contamination.*255 256---257 258## ๐ Citation259 260If you use this dataset, please cite the original paper:261 262```bibtex263@article{cipollone2025automating,264 title={Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues},265 author={Cipollone, Daniele and Izadi, Maliheh and Wang, Changjie and Scazzariello, Mariano and Ferlin, Simone and Kostiฤ, Dejan and Chiesa, Marco},266 journal={arXiv preprint arXiv:2501.05258},267 year={2025}268}269 