Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K135 likes6.6k downloads1y agoHugging Face02AlicanKiraz0 /Cybersecurity-Dataset-Fenrir-v2.1 Cybersecurity Defense Instruction-Tuning Dataset (v2.1) Created by Alican Kiraz TL;DR A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1.texttext-generation10K<n<100K151 likes3.9k downloads6mo agoHugging Face03Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K123 likes3.4k downloads1mo agoHugging Face04torchsight /cybersecurity-classification-benchmark TorchSight Cybersecurity Classification Benchmark A two-tier benchmark dataset for evaluating cybersecurity document classifiers, released with the TorchSight system. Used in: Dobrovolskyi, I. Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System. Journal of Information Security and Applications, 2026. Canonical per-model numbers live in BENCHMARK_NUMBERS.md, auto-generated from the per-prediction result JSONs… See the full description on the dataset page: https://huggingface.co/datasets/torchsight/cybersecurity-classification-benchmark.texttext-classification1K<n<10K1 likes2k downloads5mo agoHugging Face05ethanolivertroy /nist-cybersecurity-training NIST Cybersecurity Training Dataset v1.1 The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs Version 1.1 Highlights What's New in v1.1: ✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents ✅ Fixed 6,150 broken DOI links via format normalization ✅ Removed 202 malformed DOIs (double URL prefixes) ✅ Validated and fixed 124,946 total links ✅ Cataloged 72,698 broken links for future recovery ✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.texttext-generation100K<n<1M59 likes1.9k downloads1y agoHugging Face06clydeiii /cybersecurityAs found on https://raw.githubusercontent.com/aptnotes/data/master/APTnotes.json text100K<n<1M7 likes1.9k downloads3y agoHugging Face07witfoo /precinct6-cybersecurity WitFoo Precinct6 Cybersecurity Dataset Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph, mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data. Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC)… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity.text-classification1M<n<10M37 likes1.9k downloads15d agoHugging Face08rezaduty /cybersecurity-qa-v2 Cybersecurity Q&A Dataset v2 — 2.6M Examples A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics. 2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies. Statistics Source Examples Description NIST NVD CVE Database ~1,954,225 All CVEs (2002–2025): overview, severity, detection, remediation AlicanKiraz0/All-CVE-Records-Training-Dataset ~297,441 Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/rezaduty/cybersecurity-qa-v2.textquestion-answering1M<n<10M2 likes1.6k downloads4mo agoHugging Face09jcordon5 /cybersecurity-rules Cybersecurity Detection Rules Dataset This dataset contains a collection of 950 detection rules from official SIGMA, YARA, and Suricata repositories. Knowledge distillation was applied to generate questions for each rule and enrich the responses, using 0dAI-7.5B. Contents A set of detection rules for cybersecurity threat and intrusion detection in JSONL format (rules_dataset.jsonl). It contains the prompts and the associated responses. The rules have been obtained from… See the full description on the dataset page: https://huggingface.co/datasets/jcordon5/cybersecurity-rules.textn<1K10 likes1.5k downloads2y agoHugging Face10witfoo /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Version 2.1.0 (built 2026-09-22). Regenerated to address feedback from the University of Canterbury PIDS evaluation: attacks and normal traffic now share a timeline, usernames are wired into the provenance graph, mis-parsed timestamps are repaired, and every number in this card is generated from the uploaded data. Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center… See the full description on the dataset page: https://huggingface.co/datasets/witfoo/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B3 likes1.1k downloads15d agoHugging Face11Vanessasml /cybersecurity_32k_instruction_input_output Dataset Card The dataset Q&As are focused on identification of cyber threats, and text classification under the NIST taxonomy and ITC EBA IT risk classes Dataset Details Dataset Description This dataset includes a mix of public reports and news and aims to be used for cyber security risk model training. It includes 32k examples with instruction, input and output. The latter is the output from GPT. Curated by: [Vanessa Lopes] Language [EN] Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Vanessasml/cybersecurity_32k_instruction_input_output.tabular10K<n<100K20 likes1k downloads2y agoHugging Face12oi-uae /cyber-securitygated Cybersecurity Instruction-Tuning Dataset A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning, built from 198 distinct sources spanning offensive security, blue-team operations, vulnerability intelligence, cloud/AWS security, malware analysis, digital forensics, and more. Every record is normalized to the standard messages chat format and deduplicated at both file and record level. ⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.textquestion-answering1M<n<10M27 likes974 downloads25d agoHugging Face13Parsannazari12 /cybersecurity-master-dataset Cybersecurity Master Dataset Unified and deduplicated cybersecurity SFT dataset containing CTF solutions, CVE analyses, vulnerability patches, and Python coding instructions. texttext-generation100K<n<1M3 likes922 downloads1mo agoHugging Face14dpevzner /Cybersecurity_Reasoning_Dataset Cybersecurity Reasoning Dataset (Model-Agnostic) A model-agnostic re-architecture of the Cybersecurity Reasoning Dataset. The original corpus was format-bound to the Mistral/Llama ### Instruction: / ### Response: template; this dataset losslessly separates reasoning content from format, providing one neutral canonical corpus plus four per-family rendered training variants (Mistral/Llama, DeepSeek, ChatML, Gemma). Why this exists. Identical content scored 88.1 on a… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset.texttext-generationn<1K1 likes894 downloads3mo agoHugging Face15Humanlearning /CyberSecurity_OWASP-sft-dataset CyberSecurity_OWASP SFT Dataset This dataset contains verifier-gated supervised fine-tuning examples for the CyberSecurity_OWASP OpenEnv environment. Each row teaches one step of the defensive local AppSec workflow: inspect policy/code, reproduce a local authorization failure, submit a policy-tied diagnosis, patch the generated app, run visible tests, and submit the fix. Every kept trajectory is executed against the real local environment and must pass the deterministic reward… See the full description on the dataset page: https://huggingface.co/datasets/Humanlearning/CyberSecurity_OWASP-sft-dataset.text-generation0 likes859 downloads5mo agoHugging Face16Rowden /CybersecurityQAA Dataset Card for Cybersecurity Question-Answer-Assertion (QAA) Dataset Dataset Summary The Cybersecurity QAA dataset is designed to evaluate the capabilities of large language models (LLMs) in delivering cybersecurity advice and information, particularly for UK small and medium-sized enterprises (SMEs). The dataset comprises 1,563 question-answer-assertion triples across various cybersecurity topics, such as network security, data protection, and user access management.… See the full description on the dataset page: https://huggingface.co/datasets/Rowden/CybersecurityQAA.text1K<n<10K7 likes606 downloads2y agoHugging Face17WhitzardAgent /CyberSecurity-1Mgated CyberSecurity-1M A large-scale, multi-source cybersecurity knowledge dataset containing 1.19M records across 16 categories, collected exclusively for academic, non-commercial research purposes. Last updated: 2026-05-27. Disclaimer: This dataset is provided for academic research only. All content is aggregated from publicly available sources. The views, opinions, and information expressed in the dataset content do not represent the views or positions of the research team. The… See the full description on the dataset page: https://huggingface.co/datasets/WhitzardAgent/CyberSecurity-1M.texttext-generation1K<n<10K27 likes587 downloads4mo agoHugging Face18CyberNative /CyberSecurityEvalCyberNative AI for CyberSecurity Q/A Evaluation | NOT FOR TRAINING This is an evaluation dataset, please do not use for training. Tested models: CyberNative-AI/Colibri_8b_v0.1 | SCORE: 74/100 | Comments & code cognitivecomputations/dolphin-2.9-llama3-8b | SCORE: 67/100 Hermes-2-Pro-Llama-3-8B | SCORE: 65/100 segolilylabs/Lily-Cybersecurity-7B-v0.2 | SCORE: 63/100 | Comments & code cognitivecomputations/dolphin-2.9.1-llama-3-8b | FAILED TESTING (Gibberish) textn<1K22 likes586 downloads2y agoHugging Face19mariiazhiv /cybersecurity_full_question_answerstext1K<n<10K0 likes459 downloads11mo agoHugging Face20emgena /omnimcp_cybersecurity_secops_teaser 🚀 OmniMCP CyberDefense SecOps Cloud Armor (Evaluation Teaser + Turnkey MCP Server) ⚡ Official Free Community Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Production Master Package on Gumroad:👉 Purchase Full Enterprise Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout! (Starting at €249) ⚡ Activate in Cursor IDE & Claude Desktop in 30 Seconds This repository now contains a zero-dependency, turnkey Model Context Protocol (MCP)… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cybersecurity_secops_teaser.n<1K1 likes451 downloads13d agoHugging Face21Tiamz /cybersecurity-instruction-datasettext10K<n<100K0 likes438 downloads1y agoHugging Face22AlicanKiraz0 /Cybersecurity-Dataset-Heimdall-v1.1 Cybersecurity Defense Instruction-Tuning Dataset (v1.1) TL;DR 21 258 high‑quality system / user / assistant triples for training alignment‑safe, defensive‑cybersecurity LLMs. Curated from 100 000 + technical sources, rigorously cleaned and filtered to enforce strict ethical boundaries. Apache‑2.0 licensed. 1  What’s new in v1.1  (2025‑06‑21) Change v1.0 v1.1 Rows 2 500 21 258 (+760 %) Covered frameworks OWASP Top 10, NIST CSF + MITRE ATT&CK, ASD… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Heimdall-v1.1.texttext-generation10K<n<100K21 likes397 downloads1y agoHugging Face23Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes386 downloads3mo agoHugging Face24True2456 /cybersecurity-theory-sft-gemma12b Cybersecurity Theory SFT (Gemma 12B pack) Curated 21,265-row cybersecurity theory instruction pack for LoRA supervised fine-tuning. Each example is a single-turn user → assistant pair covering offensive/defensive concepts, frameworks, CTF reasoning, vulnerability catalogs, and security tooling literacy — without agent tool traces or multi-turn harness data. Paired MLX LoRA adapter trained on this pack (Nemotron 3 Super… See the full description on the dataset page: https://huggingface.co/datasets/True2456/cybersecurity-theory-sft-gemma12b.text-generation10K<n<100K2 likes374 downloads3mo agoHugging Face25theResearchNinja /benchmarkResults_violentUTF_cybersecurityBehavior Overview Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security. Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.tabular100K<n<1M1 likes329 downloads11mo agoHugging Face26Parsannazari12 /cybersecurity-sft-datasettext1M<n<10M0 likes315 downloads1mo agoHugging Face27artham123 /precinct6-cybersecurity-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/artham123/precinct6-cybersecurity-100m.tabulartext-classification100M<n<1B0 likes314 downloads4mo agoHugging Face28Cyber-security-final-project /Generated_Injected_PDFs_HARMLESS Generated Injected PDFs — HARMLESS A synthetic dataset of 1,100 PDF files built for training and evaluating structural PDF-malware detectors. It pairs benign PDFs with PDFs into which safe, non-executable "malware-shaped" objects have been injected, so a model can learn to separate the two from byte-level structure alone. ⚠️ Safety notice — read first Nothing in this dataset is real malware. Every injected payload is built from industry-standard, non-executable… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Generated_Injected_PDFs_HARMLESS.documenttabular-classification1K<n<10K0 likes306 downloads2mo agoHugging Face29ChaoticNeutrals /Cybersecurity-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing text10K<n<100K21 likes271 downloads2y agoHugging Face30mariiazhiv /cybersecurity_qa Cybersecurity QA This dataset contains instruction–response pairs focused on cybersecurity concepts.It can be used for instruction-tuned fine-tuning of LLMs Dataset Structure Format: JSONL (.jsonl) Each line is a JSON object with fields: instruction: the task or question input: optional extra context (empty string in this dataset) output: the expected answer Example: {"instruction": "What is cybersecurity's primary purpose?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/mariiazhiv/cybersecurity_qa.textquestion-answeringn<1K2 likes244 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.