Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K17 likes16k downloads4mo agoHugging Face02Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K117 likes3.4k downloads1mo agoHugging Face03yatin-superintelligence /White-Hat-Security-Agent-Prompts-600K White Hat Security Agent Prompts 600K Overview The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios. Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.texttext-generation100K<n<1M21 likes2.1k downloads7mo agoHugging Face04Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K7 likes984 downloads11mo agoHugging Face05oi-uae /cyber-securitygated Cybersecurity Instruction-Tuning Dataset A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning, built from 198 distinct sources spanning offensive security, blue-team operations, vulnerability intelligence, cloud/AWS security, malware analysis, digital forensics, and more. Every record is normalized to the standard messages chat format and deduplicated at both file and record level. ⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.textquestion-answering1M<n<10M27 likes945 downloads24d agoHugging Face06secmlr /llm-fv-security-targets LLM-FV Security Targets This dataset contains 1183 independently validated, containerized security-agent targets produced by the ucsb-mlsec/llm-fv pipelines. Contents GitHub Global Security Advisories: 394 targets OSS-Fuzz: 788 targets PoC task support: 1183 targets Exploit task support: 395 targets Patch task support: 395 targets Compressed bundle size: 58.51 GiB Vulnerability classes: {'logic_bug': 409, 'memory_vulnerability': 774} Primary languages: {'C': 201… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/llm-fv-security-targets.tabulartext-generation1K<n<10K0 likes639 downloads7h agoHugging Face07Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M1 likes563 downloads1mo agoHugging Face08heron-ai-security /stegoattack-advbench50 StegoAttack AdvBench-50 Steganographic jailbreak data generated using the StegoAttack pipeline from the paper "Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks" (Geng et al., 2025). For experiment results and analysis, see experiment.md. What is StegoAttack? StegoAttack is a jailbreak method that uses steganography to hide harmful queries inside benign-looking text. It embeds each word of a harmful query at a fixed position (e.g. the 2nd… See the full description on the dataset page: https://huggingface.co/datasets/heron-ai-security/stegoattack-advbench50.text-generationn<1K0 likes465 downloads20d agoHugging Face09adastracomputing /nixpkgs-security-patches nixpkgs-security-patches Training dataset for fine-tuning LLMs on nixpkgs security patch generation. Derived from real merged security PRs in NixOS/nixpkgs. Dataset Details 588 training examples / 66 eval examples (654 total) Format: Multi-turn tool-calling conversations in ChatML JSONL Each example is a realistic agent session: the model reads the package file, finds the upstream fix, computes hashes via tools, and submits the fix for approval Hashes and URLs… See the full description on the dataset page: https://huggingface.co/datasets/adastracomputing/nixpkgs-security-patches.texttext-generationn<1K1 likes119 downloads7mo agoHugging Face10Lots-of-LoRAs /task692_mmmlu_answer_generation_computer_security Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task692_mmmlu_answer_generation_computer_security Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task692_mmmlu_answer_generation_computer_security.texttext-generationn<1K0 likes116 downloads2y agoHugging Face11sahilempire /redsec-security-sft-v1 RedSec Security SFT v1 A chat-formatted supervised fine-tuning dataset for security-focused language models, intended for authorized red-team, penetration-testing, and defensive use. Each record is a {"messages": [...]} conversation with an optional system turn, a user turn, and an assistant turn. Split Rows train 55,459 validation 1,155 test 1,155 total 57,769 Sources and attribution This is a derivative work. It combines, reformats… See the full description on the dataset page: https://huggingface.co/datasets/sahilempire/redsec-security-sft-v1.texttext-generation10K<n<100K0 likes109 downloads2mo agoHugging Face12Lots-of-LoRAs /task733_mmmlu_answer_generation_security_studies Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task733_mmmlu_answer_generation_security_studies Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task733_mmmlu_answer_generation_security_studies.texttext-generationn<1K0 likes102 downloads2y agoHugging Face13316usman /email-security EMAIL_SECURITY A preference dataset for EMAIL_SECURITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits 80/10/10 train… See the full description on the dataset page: https://huggingface.co/datasets/316usman/email-security.texttext-generation1K<n<10K0 likes102 downloads8d agoHugging Face14Lateos /offensive_security_dataset Preview samples from the Lateos Red Team training pipeline. Contains a deterministic 3% sample (seed 42) of the full corpus, drawn from the exact production datasets used to train red-team and offensive-security models. These samples let you evaluate schema, quality, and provenance before licensing the full datasets. All data is dual-use security research: attack playbooks, vulnerability analysis, tool usage guides, and decision frameworks for authorized penetration testing.… See the full description on the dataset page: https://huggingface.co/datasets/Lateos/offensive_security_dataset.text-generation0 likes86 downloads2mo agoHugging Face15Publicus /cvefixes-security-autoformal-span-cache CVEfixes Security source spans This dataset contains exact, source-bound prose, code, and diff spans derived from the original-data configuration of Publicus/cvefixes-security-ir-graphrag at revision 6fd5918bed34f8851430e74a149502587a953fe2. The underlying source is hitoshura25/cvefixes at revision d4f5c4ea65329d9ccbb8a3b3149e5d06eda5edb2. The extraction considered all 12,987 original rows across all three shards. It retained 9,402 source rows and excluded 3,585. The release… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-autoformal-span-cache.tabulartext-generation100K<n<1M0 likes82 downloads5d agoHugging Face16emgena /omnimcp_supabase_row_level_security_ai_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_supabase_row_level_security_ai_teaser.texttext-generationn<1K0 likes72 downloads19d agoHugging Face17davidfoss /bitcoin-security-reasoning-100k Dataset Card for Bitcoin Security Reasoning 100K 100,000 high-quality synthetic training samples for fine-tuning LLMs on Bitcoin protocol security analysis. Teaches models to analyze vulnerability clusters, form security hypotheses, and generate differential testing code. Dataset Details Dataset Description This dataset contains structured security reasoning chains for Bitcoin protocol vulnerabilities. Each sample presents a cluster of causal… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/bitcoin-security-reasoning-100k.texttext-generation100K<n<1M0 likes71 downloads9mo agoHugging Face18Dhanjo /ai-agent-security-dataset AI Agent Security and System Prompt Leakage Dataset Dataset Overview This dataset was created for research on AI agent security, with a specific focus on system prompt leakage, jailbreak resistance, and security-aligned fine-tuning. The dataset evaluates how often AI agents reveal confidential information embedded inside their system prompts when exposed to adversarial prompts. It also compares the behavior of a baseline language model against a model fine-tuned using… See the full description on the dataset page: https://huggingface.co/datasets/Dhanjo/ai-agent-security-dataset.tabulartext-generation1K<n<10K0 likes56 downloads5mo agoHugging Face19sublime-security /mql-benchmark MQL Benchmark A benchmark for evaluating natural language → MQL (Message Query Language) generation. MQL is a DSL used at Sublime Security for email threat detection. Dataset Summary Split Examples Purpose train 21,654 Few-shot examples and fine-tuning validation 4,650 Prompt / hyperparameter tuning test 4,326 Final evaluation — use sparingly Total: 30,630 examples across four difficulty tiers and four prompt styles. Each example is a (nl_prompt… See the full description on the dataset page: https://huggingface.co/datasets/sublime-security/mql-benchmark.text-generation10K<n<100K0 likes54 downloads5mo agoHugging Face20Neura-parse /quantum-cryptography-and-post-quantum-security Neura Parse — Quantum Cryptography & Post-Quantum Security A deep vertical on cryptography that uses quantum mechanics and on classical cryptography built to resist quantum attack. It covers quantum key distribution (BB84, B92, six-state, SARG04, E91, BBM92, decoy-state, MDI-QKD, TF-QKD, CV-QKD), device-independent protocols, composable and finite-key security proofs, quantum hacking with countermeasures, classical post-processing (reconciliation, privacy amplification… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-cryptography-and-post-quantum-security.tabularquestion-answering100K<n<1M0 likes52 downloads3mo agoHugging Face21Arno-MHL /ios-security-vulnerabilities-swift-objc iOS Security Vulnerabilities Dataset (Swift & Objective-C) A comprehensive dataset of 27 real-world iOS security vulnerability patterns in Swift and Objective-C, covering all OWASP Mobile Top 10 (2024) categories with vulnerable code, secure fixes, attack scenarios, and detection guidance. 🎯 Purpose This is the first dedicated iOS/Swift/Objective-C security vulnerability dataset on Hugging Face. While existing datasets (TitanVul, DiverseVul, CleanVul) focus on… See the full description on the dataset page: https://huggingface.co/datasets/Arno-MHL/ios-security-vulnerabilities-swift-objc.texttext-generationn<1K3 likes45 downloads6mo agoHugging Face22thuaannn /prewise-security-adapter-training-v5 Prewise Security Adapter Training V5 Bộ dữ liệu và công cụ Kaggle hoàn chỉnh để fine-tune ba LoRA trên cùng base model Qwen/Qwen3.5-4B: message-context-adapter web-context-adapter explanation-adapter phone-intelligence là HTTP provider bên ngoài, không phải LoRA. Packager tạo entry này ở trạng thái enabled: false để backend vẫn có đủ bốn runtime contract. Quy mô Adapter Train Validation Test Tổng Message Context 32.767 3.196 3.992 39.955 Web Context… See the full description on the dataset page: https://huggingface.co/datasets/thuaannn/prewise-security-adapter-training-v5.documenttext-generation10K<n<100K0 likes45 downloads3mo agoHugging Face23tuandunghcmut /combine-llm-security-benchmarkgated Combined LLM Security Benchmark 🔐 A comprehensive, unified benchmark dataset for evaluating Large Language Models (LLMs) on cybersecurity tasks. This dataset combines 10 security benchmarks into a standardized format with 18,059 examples across 5 task types. 📊 Dataset Summary This dataset consolidates multiple security-focused benchmarks into a single, easy-to-use format for comprehensive LLM evaluation across various cybersecurity domains: Total Examples: 18,059 Total… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/combine-llm-security-benchmark.textquestion-answering10K<n<100K6 likes44 downloads1y agoHugging Face24guinansu /paragen-security-sft-alpaca paragen-security-sft-alpaca Alpaca-format instruction-tuning data used to train the Vanilla security baseline (and as the source for the tokenized multi-stream cache used to train the Stream(Ours) security checkpoint) in the paragen_llm security/prompt-injection-robustness experiments (Table 3: TensorTrust, Gandalf, Purple, RuLES, StruQ, NESSiE, IFEval). Format: JSONL, one object per line, fields instruction / input / output (standard Alpaca schema). Size: 48,538 examples.… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/paragen-security-sft-alpaca.texttext-generation10K<n<100K0 likes43 downloads22d agoHugging Face25z0n3x /oak-security-sft OAK (On-chain Attack Knowledge) Security SFT Corpus A bilingual (EN+RU) instruction-tuning dataset for training language models to analyze on-chain attacks, DeFi exploits, blockchain forensics, and crypto security incidents. Every answer is grounded in the OAK taxonomy v0.1 — a structured knowledge base of adversary tactics, techniques, mitigations, software tools, threat groups, and real-world on-chain incidents. Dataset Overview Metric Value Total… See the full description on the dataset page: https://huggingface.co/datasets/z0n3x/oak-security-sft.question-answering1K<n<10K2 likes39 downloads4mo agoHugging Face26immu4989 /dspy-security-bench-trainset-workspace dspy-security-bench: workspace trainset (v0.1) This is the synthetic, environment-grounded query-only trainset used to optimize DSPy programs in v0.1 of dspy-security-bench, a benchmark that measures whether DSPy prompt optimization affects the prompt-injection robustness of agentic LLM programs. What's in here 192 query / ground-truth pairs grounded in the AgentDojo workspace suite's default environment (calendar, inbox, files). {"prompt": "What is the… See the full description on the dataset page: https://huggingface.co/datasets/immu4989/dspy-security-bench-trainset-workspace.textquestion-answeringn<1K1 likes39 downloads3mo agoHugging Face27logicBombExe /turkish_cyber_security_controls_dataset Turkish Cyber Security Controls Dataset Veri Kümesi Özeti Bu veri kümesi; siber güvenlik kontrolleri, kontrol seçimi ve güvenli mimari tasarımı hakkında hazırlanmış 800 Türkçe kullanıcı-asistan konuşma çifti içerir. Toplam 1.600 mesajdan oluşan koleksiyon, Türkçe siber güvenlik soru-cevap ve instruction-tuning çalışmalarını desteklemek amacıyla hazırlanmıştır. İçerik geliştirilirken başta NIST SP 800-53 Rev. 5 kontrol kataloğu olmak üzere risk temelli kontrol… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_dataset.texttext-generationn<1K3 likes38 downloads3mo agoHugging Face28huzaifas-sidhpurwala /RedHat-security-VeX Dataset Card for RedHat-security-VeX This Dataset is extracted from publicly available Vulnerability Exploitability eXchange (VEX) files published by Red Hat. Dataset Details Red Hat security data is a central source of truth for Red Hat products regarding published, known vulnerabilities. This data is published in form of Vulnerability Exploitability eXchange (VEX) available at: https://security.access.redhat.com/data/csaf/v2/vex/ This Dataset is created by extracting… See the full description on the dataset page: https://huggingface.co/datasets/huzaifas-sidhpurwala/RedHat-security-VeX.textfeature-extraction10K<n<100K9 likes37 downloads7mo agoHugging Face29sumitguha13 /ai-agent-security-sft-dpo AI Agent Security — SFT + DPO Fine-tuning data for teaching an AI agent to protect its confidential configuration without becoming uselessly over-cautious. Built for thesreedath/gemma-2-2b-qa-sft and derived from Dhanjo/ai-agent-security-dataset. Why the helpfulness axis exists leakage_score in the source dataset is one-sided: a model that refuses every request scores a perfect 0.0. An existing fine-tune reported 0.0114 mean leakage (down from 0.4611 baseline)… See the full description on the dataset page: https://huggingface.co/datasets/sumitguha13/ai-agent-security-sft-dpo.tabulartext-generation10K<n<100K0 likes36 downloads1mo agoHugging Face30davidquicast /information-security-policies-qa-distiset Dataset Card for information-security-policies-qa-distiset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/daqc/information-security-policies-qa-distiset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/davidquicast/information-security-policies-qa-distiset.tabulartext-generationn<1K0 likes33 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.