Bouquets/Cybersecurity-LLM-CVE
2025.06.07 Updated data code :https://github.com/Bouquets-ai/Data-Processing/blob/main/CVE-Data.py Change 121 lines of code (keyword="CVE-2025") to obtain the required CVE time 93 lines of code (json record=) to change the required format Cybersecurity-LLM-CVE Dataset Introduction ๐ Overview ๐ก๏ธ An open-source cybersecurity vulnerability dataset designed for training/evaluating Large Language Models (LLMs) in security domains. Covers all public CVE IDs from January 1, 2021 to April 9, 2025โฆ See the full description on the dataset page: https://huggingface.co/datasets/Bouquets/Cybersecurity-LLM-CVE.
2025.06.07 Updated data code :https://github.com/Bouquets-ai/Data-Processing/blob/main/CVE-Data.py
Change 121 lines of code (keyword="CVE-2025") to obtain the required CVE time
93 lines of code (json record=) to change the required format
Cybersecurity-LLM-CVE Dataset Introduction ๐
Overview ๐ก๏ธ
An open-source cybersecurity vulnerability dataset designed for training/evaluating Large Language Models (LLMs) in security domains. Covers all public CVE IDs from January 1, 2021 to April 9, 2025, with structured data including vulnerability details (affected products) and disclosure dates.
Key Highlights โจ
Comprehensive ๐: 5-year global vulnerability records (software/hardware/protocols)
Authoritative Integration ๐: Data sourced from NVD, CVE lists, and community contributions, double-verified
Usage Declaration ๐ก
Typical Use Cases ๐ ๏ธ
๐ง LLM-powered vulnerability patch generation
๐ค Automated threat intelligence analysis
๐ง Penetration testing toolkit enhancement
Final Notes ๐
๐ฌ Feel free to leave suggestions in the community forum! I'll keep optimizing based on your feedback~ ๐ฅฐ
