Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K135 likes5k downloads1y agoHugging Face02AlicanKiraz0 /Cybersecurity-Dataset-Fenrir-v2.1 Cybersecurity Defense Instruction-Tuning Dataset (v2.1) Created by Alican Kiraz TL;DR A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1.texttext-generation10K<n<100K154 likes3.6k downloads6mo agoHugging Face03Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K133 likes2.3k downloads1mo agoHugging Face04jcordon5 /cybersecurity-rules Cybersecurity Detection Rules Dataset This dataset contains a collection of 950 detection rules from official SIGMA, YARA, and Suricata repositories. Knowledge distillation was applied to generate questions for each rule and enrich the responses, using 0dAI-7.5B. Contents A set of detection rules for cybersecurity threat and intrusion detection in JSONL format (rules_dataset.jsonl). It contains the prompts and the associated responses. The rules have been obtained from… See the full description on the dataset page: https://huggingface.co/datasets/jcordon5/cybersecurity-rules.textn<1K10 likes1.5k downloads2y agoHugging Face05dpevzner /Cybersecurity_Reasoning_Dataset Cybersecurity Reasoning Dataset (Model-Agnostic) A model-agnostic re-architecture of the Cybersecurity Reasoning Dataset. The original corpus was format-bound to the Mistral/Llama ### Instruction: / ### Response: template; this dataset losslessly separates reasoning content from format, providing one neutral canonical corpus plus four per-family rendered training variants (Mistral/Llama, DeepSeek, ChatML, Gemma). Why this exists. Identical content scored 88.1 on a… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset.texttext-generationn<1K1 likes823 downloads3mo agoHugging Face06CyberNative /CyberSecurityEval Evaluation caution — 2026-10-07 This dataset has unresolved provenance and reliability defects. Its question sources, generation process, original contributors, validation records, and dataset license have not been established. The original question file is preserved unchanged. No dataset license is asserted by this correction. A text audit of the 500-row file found 372 items with exactly four A–D choices and a matching parsed key. Their keys are A: 57, B: 70, C: 103, D: 142.… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/CyberSecurityEval.textn<1K23 likes487 downloads4d agoHugging Face07Rowden /CybersecurityQAA Dataset Card for Cybersecurity Question-Answer-Assertion (QAA) Dataset Dataset Summary The Cybersecurity QAA dataset is designed to evaluate the capabilities of large language models (LLMs) in delivering cybersecurity advice and information, particularly for UK small and medium-sized enterprises (SMEs). The dataset comprises 1,563 question-answer-assertion triples across various cybersecurity topics, such as network security, data protection, and user access management.… See the full description on the dataset page: https://huggingface.co/datasets/Rowden/CybersecurityQAA.text1K<n<10K7 likes445 downloads2y agoHugging Face08WhitzardAgent /CyberSecurity-1Mgated CyberSecurity-1M A large-scale, multi-source cybersecurity knowledge dataset containing 1.19M records across 16 categories, collected exclusively for academic, non-commercial research purposes. Last updated: 2026-05-27. Disclaimer: This dataset is provided for academic research only. All content is aggregated from publicly available sources. The views, opinions, and information expressed in the dataset content do not represent the views or positions of the research team. The… See the full description on the dataset page: https://huggingface.co/datasets/WhitzardAgent/CyberSecurity-1M.texttext-generation1K<n<10K28 likes443 downloads5mo agoHugging Face09AlicanKiraz0 /Cybersecurity-Dataset-Heimdall-v1.1 Cybersecurity Defense Instruction-Tuning Dataset (v1.1) TL;DR 21 258 high‑quality system / user / assistant triples for training alignment‑safe, defensive‑cybersecurity LLMs. Curated from 100 000 + technical sources, rigorously cleaned and filtered to enforce strict ethical boundaries. Apache‑2.0 licensed. 1  What’s new in v1.1  (2025‑06‑21) Change v1.0 v1.1 Rows 2 500 21 258 (+760 %) Covered frameworks OWASP Top 10, NIST CSF + MITRE ATT&CK, ASD… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Heimdall-v1.1.texttext-generation10K<n<100K21 likes395 downloads1y agoHugging Face10mariiazhiv /cybersecurity_full_question_answerstext1K<n<10K0 likes287 downloads1y agoHugging Face11ChaoticNeutrals /Cybersecurity-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing text10K<n<100K21 likes216 downloads2y agoHugging Face12mariiazhiv /cybersecurity_qa Cybersecurity QA This dataset contains instruction–response pairs focused on cybersecurity concepts.It can be used for instruction-tuned fine-tuning of LLMs Dataset Structure Format: JSONL (.jsonl) Each line is a JSON object with fields: instruction: the task or question input: optional extra context (empty string in this dataset) output: the expected answer Example: {"instruction": "What is cybersecurity's primary purpose?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/mariiazhiv/cybersecurity_qa.textquestion-answeringn<1K2 likes203 downloads1y agoHugging Face13hcnote /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanitytool, with only data scoring 4.5 or… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-High-Quality-Dataset.text100K<n<1M11 likes172 downloads8mo agoHugging Face14stindardlogic /cybersecurity-sft-100k Cybersecurity SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering. Dataset Description This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.texttext-generation100K<n<1M0 likes167 downloads3mo agoHugging Face15Bouquets /Cybersecurity-LLM-CVE2025.06.07 Updated data code :https://github.com/Bouquets-ai/Data-Processing/blob/main/CVE-Data.py Change 121 lines of code (keyword="CVE-2025") to obtain the required CVE time 93 lines of code (json record=) to change the required format Cybersecurity-LLM-CVE Dataset Introduction 🚀 Overview 🛡️ An open-source cybersecurity vulnerability dataset designed for training/evaluating Large Language Models (LLMs) in security domains. Covers all public CVE IDs from January 1, 2021 to April 9, 2025… See the full description on the dataset page: https://huggingface.co/datasets/Bouquets/Cybersecurity-LLM-CVE.text100K<n<1M16 likes153 downloads1y agoHugging Face16atmike /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanity tool, with only data scoring 4.5… See the full description on the dataset page: https://huggingface.co/datasets/atmike/Cybersecurity-High-Quality-Dataset.text100K<n<1M0 likes108 downloads4mo agoHugging Face17Chemically-motivated /CyberSecurityDataset Dataset Card for Cyber Security Dataset This dataset provides a collection of curated data points related to cybersecurity, focusing on penetration testing, known exploits, and vulnerability analysis. It is intended to aid researchers, educators, and developers in building AI tools for cybersecurity applications. Dataset Details Dataset Description This dataset contains labeled information about exploits, vulnerabilities, and penetration testing techniques.… See the full description on the dataset page: https://huggingface.co/datasets/Chemically-motivated/CyberSecurityDataset.textn<1K4 likes92 downloads2y agoHugging Face18ScoutieAutoML /cybersecurity_news_telegram_dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic Cybersecurity, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the channel. views -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/cybersecurity_news_telegram_dataset.tabulartext-classification10K<n<100K3 likes83 downloads2y agoHugging Face19ansulev /cybersecurity-dataset-fenrir Cybersecurity Defense Instruction-Tuning Dataset (v2.1) Created by Alican Kiraz TL;DR A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/cybersecurity-dataset-fenrir.texttext-generation10K<n<100K1 likes70 downloads6mo agoHugging Face20Bouquets /DeepSeek-V3-Distill-Cybersecurity-en 🔐 CyberSecPentest-DS: Deepseek-V3 Distilled Dataset 📂 Dataset Overview:This is a high-quality distilled dataset 🧪, specialized in the cybersecurity penetration testing domain, generated using Deepseek-V3 🚀. It provides curated knowledge for red-teaming, vulnerability assessment, and ethical hacking research. 💡 Key Features: 🛡️ Focus: Penetration testing, exploit development, and security assessments. 🧠 Source: Knowledge distilled from Deepseek-V3 for accuracy &… See the full description on the dataset page: https://huggingface.co/datasets/Bouquets/DeepSeek-V3-Distill-Cybersecurity-en.text1K<n<10K0 likes68 downloads1y agoHugging Face21ystemsrx /Cybersecurity-ShareGPT-ChineseEnglish 网络安全中文数据集 (ShareGPT 格式) 本数据集是一个关于网络安全的中文对话数据集,采用 ShareGPT 格式,适用于语言模型的训练和微调。该数据集包含多个与网络安全相关的对话,能够帮助语言模型在网络安全领域进行学习与优化。数据集以 json 和 jsonl 两种格式提供,便于用户灵活使用。 数据集内容 该数据集以网络安全为主题的对话数据为核心,旨在用于以下任务: 语言模型的训练与微调 对话生成任务 网络安全相关的对话系统构建 研究网络安全领域的自动化问答系统 数据格式 每个数据样本的格式遵循 ShareGPT 的对话格式,结构如下: { "conversations": [ { "from": "system", "value": "..." }, { "from": "human", "value": "..." }… See the full description on the dataset page: https://huggingface.co/datasets/ystemsrx/Cybersecurity-ShareGPT-Chinese.text10K<n<100K22 likes63 downloads2y agoHugging Face22Mohabahmed03 /Alpaca_Dataset_CyberSecurity_Smaller_2.0text10K<n<100K0 likes58 downloads1y agoHugging Face23hcnote /Cybersecurity-bigDataset 🛡️ CyberSec-MegaDataset v1.0 全球首个开源超大规模网络安全数据集 | 新疆幻城网安科技有限公司 | 截止至 2026-01-01 🔍 核心价值:为什么它是全球最大的开源网络安全数据集? CyberSec-MegaDataset 由 新疆幻城网安科技有限公司 于 2026-01-01 完成最终整合与验证,是目前全球规模最大、覆盖最全、质量最高的开源网络安全数据集合。本数据集严格覆盖截止至 2026 年 1 月 1 日的主流开源安全数据源,并融合公司自研的千万级真实攻防日志与代码样本,经 MinHash + SimHash 双重去重、全链路隐私脱敏(符合 GDPR/CCPA 标准)与 专家级语义标注,专为训练 企业级本地化 AI 安全大模型(如 DeepSeek-14B、Qwen-Code-30B MOE)而设计。 ✅ 关键突破: **覆盖度 100%**:整合市面上 99.8% 的主流开源网络安全数据源(含 32 个权威平台),填补了企业内网安全训练数据的全球空白。 规模… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-bigDataset.text100K<n<1M4 likes56 downloads9mo agoHugging Face24humanwill-ai /cybersecurity-initial424 HumanWill Cybersecurity Benchmark Measuring harmful refusals in AI models. 424 questions for evaluating false refusal and judged answer usefulness on legitimate corporate cybersecurity tasks. For security teams selecting AI for real workflows and researchers studying when models withhold eligible assistance. Framework and CLI on GitHub · Interactive results · Full report (PDF) · HumanWill What is included Topic Questions Vulnerability validation 81… See the full description on the dataset page: https://huggingface.co/datasets/humanwill-ai/cybersecurity-initial424.texttext-generationn<1K0 likes54 downloads20d agoHugging Face25safouene99999 /Cybersecurity_QAtext10K<n<100K0 likes52 downloads1y agoHugging Face26ArkhAngelLifeJiggy /CyberSecurityDataset Dataset Card for Cyber Security Dataset This dataset provides a collection of curated data points related to cybersecurity, focusing on penetration testing, known exploits, and vulnerability analysis. It is intended to aid researchers, educators, and developers in building AI tools for cybersecurity applications. Dataset Details Dataset Description This dataset contains labeled information about exploits, vulnerabilities, and penetration testing… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/CyberSecurityDataset.textn<1K0 likes52 downloads25d agoHugging Face27deer-sec /deer_sec-japanese-cybersecurity-chatml-v2 deer_sec-japanese-cybersecurity-chatml-v2.0 📊 Dataset Details Total Rows: 83,562 件 File Size: 約 492.5 MB Format: JSONL (ChatML形式) Language: 日本語 (Japanese) 概要 (Overview) 本データセットは、高度なサイバー防御と脅威インテリジェンスに特化したインストラクション・チューニング用のデータセットです。AlicanKiraz0様によって公開された Cybersecurity-Dataset-Fenrir-v2.0 を元に構築されています。 元の膨大な英語データの中から約 99.5% (83,562件) を抽出し、翻訳特化モデルである translategemma:12b を用いて高品質な日本語へ翻訳しました。その後、LLMのファインチューニング(LoRA等)にそのまま利用できるよう、厳格なデータクレンジングと整形を行っています。 特徴… See the full description on the dataset page: https://huggingface.co/datasets/deer-sec/deer_sec-japanese-cybersecurity-chatml-v2.texttext-generation10K<n<100K0 likes49 downloads5mo agoHugging Face28ChipHolmes /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.texttext-generation10K<n<100K1 likes48 downloads3mo agoHugging Face29Sidhantyadav /cybersecurity_reasoning_thinking_jsontextn<1K0 likes48 downloads20d agoHugging Face30hcnote /High-quality-cybersecurity-datasets 网络安全高质量数据集 📊 数据集概述 这是一个经过严格筛选和清洗的高质量网络安全领域数据集,共包含 277,707 条优质数据记录。数据集通过 AI 标注和人工评分双重筛选机制,确保了内容的专业性、准确性和实用价值。 📁 数据格式 数据集采用 JSONL 格式(每行一个独立的 JSON 对象),结构如下: { "instruction": "问题/指令/提示词", "output": "回答/输出/响应", "input": "输入内容(可选字段)", "id": "唯一标识符(UUID格式)" } 🎯 数据集特点 1. 高质量保证 ✅ 经过 AI 自动标注 ✅ 人工专家评分审核 ✅ 多轮过滤筛选机制 ✅ 去除低质量和重复数据 2. 内容全面丰富 涵盖网络安全的各个主要领域,从理论到实践,从防御到攻击 3. 多语言支持 中英文混合内容 涵盖国内外最新安全动态 4. 实用性强… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/High-quality-cybersecurity-datasets.text100K<n<1M0 likes46 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.