Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K135 likes5k downloads1y agoHugging Face02AlicanKiraz0 /Cybersecurity-Dataset-Fenrir-v2.1 Cybersecurity Defense Instruction-Tuning Dataset (v2.1) Created by Alican Kiraz TL;DR A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1.texttext-generation10K<n<100K153 likes3.6k downloads6mo agoHugging Face03Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K132 likes2.3k downloads1mo agoHugging Face04torchsight /cybersecurity-classification-benchmark TorchSight Cybersecurity Classification Benchmark A two-tier benchmark dataset for evaluating cybersecurity document classifiers, released with the TorchSight system. Used in: Dobrovolskyi, I. Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System. Journal of Information Security and Applications, 2026. Canonical per-model numbers live in BENCHMARK_NUMBERS.md, auto-generated from the per-prediction result JSONs… See the full description on the dataset page: https://huggingface.co/datasets/torchsight/cybersecurity-classification-benchmark.texttext-classification1K<n<10K1 likes1.8k downloads5mo agoHugging Face05ethanolivertroy /nist-cybersecurity-training NIST Cybersecurity Training Dataset v1.1 The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs Version 1.1 Highlights What's New in v1.1: ✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents ✅ Fixed 6,150 broken DOI links via format normalization ✅ Removed 202 malformed DOIs (double URL prefixes) ✅ Validated and fixed 124,946 total links ✅ Cataloged 72,698 broken links for future recovery ✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.texttext-generation100K<n<1M59 likes1.6k downloads1y agoHugging Face06clydeiii /cybersecurityAs found on https://raw.githubusercontent.com/aptnotes/data/master/APTnotes.json text100K<n<1M7 likes1.6k downloads3y agoHugging Face07jcordon5 /cybersecurity-rules Cybersecurity Detection Rules Dataset This dataset contains a collection of 950 detection rules from official SIGMA, YARA, and Suricata repositories. Knowledge distillation was applied to generate questions for each rule and enrich the responses, using 0dAI-7.5B. Contents A set of detection rules for cybersecurity threat and intrusion detection in JSONL format (rules_dataset.jsonl). It contains the prompts and the associated responses. The rules have been obtained from… See the full description on the dataset page: https://huggingface.co/datasets/jcordon5/cybersecurity-rules.textn<1K10 likes1.5k downloads2y agoHugging Face08witfoo /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph, mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data. Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B3 likes1.2k downloads18d agoHugging Face09oi-uae /cyber-securitygated Cybersecurity Instruction-Tuning Dataset A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning, built from 198 distinct sources spanning offensive security, blue-team operations, vulnerability intelligence, cloud/AWS security, malware analysis, digital forensics, and more. Every record is normalized to the standard messages chat format and deduplicated at both file and record level. ⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.textquestion-answering1M<n<10M29 likes1.1k downloads28d agoHugging Face10Parsannazari12 /cybersecurity-master-dataset Cybersecurity Master Dataset Unified and deduplicated cybersecurity SFT dataset containing CTF solutions, CVE analyses, vulnerability patches, and Python coding instructions. texttext-generation100K<n<1M3 likes841 downloads1mo agoHugging Face11Vanessasml /cybersecurity_32k_instruction_input_output Dataset Card The dataset Q&As are focused on identification of cyber threats, and text classification under the NIST taxonomy and ITC EBA IT risk classes Dataset Details Dataset Description This dataset includes a mix of public reports and news and aims to be used for cyber security risk model training. It includes 32k examples with instruction, input and output. The latter is the output from GPT. Curated by: [Vanessa Lopes] Language [EN] Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Vanessasml/cybersecurity_32k_instruction_input_output.tabular10K<n<100K20 likes838 downloads2y agoHugging Face12dpevzner /Cybersecurity_Reasoning_Dataset Cybersecurity Reasoning Dataset (Model-Agnostic) A model-agnostic re-architecture of the Cybersecurity Reasoning Dataset. The original corpus was format-bound to the Mistral/Llama ### Instruction: / ### Response: template; this dataset losslessly separates reasoning content from format, providing one neutral canonical corpus plus four per-family rendered training variants (Mistral/Llama, DeepSeek, ChatML, Gemma). Why this exists. Identical content scored 88.1 on a… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset.texttext-generationn<1K1 likes823 downloads3mo agoHugging Face13rezaduty /cybersecurity-qa-v2 Cybersecurity Q&A Dataset v2 — 2.6M Examples A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics. 2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies. Statistics Source Examples Description NIST NVD CVE Database ~1,954,225 All CVEs (2002–2025): overview, severity, detection, remediation AlicanKiraz0/All-CVE-Records-Training-Dataset ~297,441 Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/rezaduty/cybersecurity-qa-v2.textquestion-answering1M<n<10M3 likes660 downloads4mo agoHugging Face14Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes512 downloads3mo agoHugging Face15CyberNative /CyberSecurityEval Evaluation caution — 2026-10-07 This dataset has unresolved provenance and reliability defects. Its question sources, generation process, original contributors, validation records, and dataset license have not been established. The original question file is preserved unchanged. No dataset license is asserted by this correction. A text audit of the 500-row file found 372 items with exactly four A–D choices and a matching parsed key. Their keys are A: 57, B: 70, C: 103, D: 142.… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/CyberSecurityEval.textn<1K22 likes487 downloads4d agoHugging Face16Rowden /CybersecurityQAA Dataset Card for Cybersecurity Question-Answer-Assertion (QAA) Dataset Dataset Summary The Cybersecurity QAA dataset is designed to evaluate the capabilities of large language models (LLMs) in delivering cybersecurity advice and information, particularly for UK small and medium-sized enterprises (SMEs). The dataset comprises 1,563 question-answer-assertion triples across various cybersecurity topics, such as network security, data protection, and user access management.… See the full description on the dataset page: https://huggingface.co/datasets/Rowden/CybersecurityQAA.text1K<n<10K7 likes445 downloads2y agoHugging Face17WhitzardAgent /CyberSecurity-1Mgated CyberSecurity-1M A large-scale, multi-source cybersecurity knowledge dataset containing 1.19M records across 16 categories, collected exclusively for academic, non-commercial research purposes. Last updated: 2026-05-27. Disclaimer: This dataset is provided for academic research only. All content is aggregated from publicly available sources. The views, opinions, and information expressed in the dataset content do not represent the views or positions of the research team. The… See the full description on the dataset page: https://huggingface.co/datasets/WhitzardAgent/CyberSecurity-1M.texttext-generation1K<n<10K27 likes443 downloads5mo agoHugging Face18AlicanKiraz0 /Cybersecurity-Dataset-Heimdall-v1.1 Cybersecurity Defense Instruction-Tuning Dataset (v1.1) TL;DR 21 258 high‑quality system / user / assistant triples for training alignment‑safe, defensive‑cybersecurity LLMs. Curated from 100 000 + technical sources, rigorously cleaned and filtered to enforce strict ethical boundaries. Apache‑2.0 licensed. 1  What’s new in v1.1  (2025‑06‑21) Change v1.0 v1.1 Rows 2 500 21 258 (+760 %) Covered frameworks OWASP Top 10, NIST CSF + MITRE ATT&CK, ASD… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Heimdall-v1.1.texttext-generation10K<n<100K21 likes395 downloads1y agoHugging Face19Tiamz /cybersecurity-instruction-datasettext10K<n<100K0 likes391 downloads2y agoHugging Face20artham123 /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B0 likes317 downloads4mo agoHugging Face21Parsannazari12 /cybersecurity-sft-datasettext1M<n<10M0 likes316 downloads1mo agoHugging Face22mariiazhiv /cybersecurity_full_question_answerstext1K<n<10K0 likes287 downloads1y agoHugging Face23theResearchNinja /benchmarkResults_violentUTF_cybersecurityBehavior Overview Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security. Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.tabular100K<n<1M1 likes285 downloads11mo agoHugging Face24ChaoticNeutrals /Cybersecurity-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing text10K<n<100K21 likes216 downloads2y agoHugging Face25AlicanKiraz0 /Cybersecurity-Dataset-v1 Cybersecurity Defense Training Dataset Dataset Description This dataset contains 2,500 high-quality instruction-response pairs focused on defensive cybersecurity education. The dataset is designed to train AI models to provide accurate, detailed, and ethically-aligned guidance on information security principles while refusing to assist with malicious activities. Dataset Summary Language: English License: Apache 2.0 Format: Parquet Size: 2,500 rows Domain:… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-v1.texttext-classification1K<n<10K16 likes207 downloads1y agoHugging Face26mariiazhiv /cybersecurity_qa Cybersecurity QA This dataset contains instruction–response pairs focused on cybersecurity concepts.It can be used for instruction-tuned fine-tuning of LLMs Dataset Structure Format: JSONL (.jsonl) Each line is a JSON object with fields: instruction: the task or question input: optional extra context (empty string in this dataset) output: the expected answer Example: {"instruction": "What is cybersecurity's primary purpose?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/mariiazhiv/cybersecurity_qa.textquestion-answeringn<1K2 likes203 downloads1y agoHugging Face27whybe-choi /kovidore-v2-cybersecurity-beirKoViDoRe v2 : Cybersecurity This dataset, Cybersecurity, is a corpus of technical reports on cyber threat trends and security incident responses in Korea, intended for complex-document understanding tasks. It is one of the 4 corpora comprising the KoViDoRe v2 Benchmark. Links Github: https://github.com/whybe-choi/kovidore-benchmark Collection: https://huggingface.co/collections/whybe-choi/kovidore-benchmark-beir-v2 Data Generation Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/kovidore-v2-cybersecurity-beir.documentdocument-question-answering1K<n<10K1 likes195 downloads9mo agoHugging Face28beatsprom /cybersecurity-soc-threat-hunting-sft-dpo-2026 🛡️ Enterprise Cybersecurity AI, SOC Tier-3 & Threat Hunting SFT/DPO Dataset (2026) High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SOC Tier-3 Chain-of-Thought (<thought>) kill-chain diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior SOC Threat Hunters, Incident Responders, and Red-Team Defense Architects. 📊 Dataset Architecture & Highlights… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cybersecurity-soc-threat-hunting-sft-dpo-2026.texttext-generationn<1K2 likes194 downloads1mo agoHugging Face29hcnote /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanitytool, with only data scoring 4.5 or… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-High-Quality-Dataset.text100K<n<1M11 likes172 downloads8mo agoHugging Face30stindardlogic /cybersecurity-sft-100k Cybersecurity SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering. Dataset Description This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.texttext-generation100K<n<1M0 likes167 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.