Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K18 likes17k downloads4mo agoHugging Face02Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K133 likes2.3k downloads1mo agoHugging Face03gussieIsASuccessfulWarlock /security_instruct_mcq_2481textn<1K0 likes2.2k downloads2y agoHugging Face04CyberNative /Code_Vulnerability_Security_DPO Secure-code dataset: correction and audit notice (2026-10-06) The chosen and rejected fields are original model-generated preference labels. They are not verified secure/insecure classifications. Do not use these labels as security ground truth or use the dataset as a validated secure-code training or evaluation set without your own contextual validation. The historical description below overstates security, optimization, realism and validation. Those claims are superseded by… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.text1K<n<10K173 likes1.3k downloads4d agoHugging Face05secmlr /llm-fv-security-targets LLM-FV Security Targets This dataset contains 889 independently validated, containerized security-agent targets produced by the ucsb-mlsec/llm-fv pipelines. Contents GitHub Global Security Advisories: 446 targets OSS-Fuzz: 442 targets PoC task support: 889 targets Exploit task support: 447 targets Patch task support: 447 targets Compressed bundle size: 42.41 GiB Vulnerability classes: {'logic_bug': 450, 'memory_vulnerability': 439} Primary languages: {'C': 118… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/llm-fv-security-targets.tabulartext-generationn<1K0 likes1.1k downloads2d agoHugging Face06pAILabs /infosec-security-qatext10K<n<100K12 likes651 downloads2y agoHugging Face07OpenClaw /clawhub-security-signals ClawHub Security Signals 🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale. This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree. Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.tabulartext-classification10K<n<100K53 likes343 downloads4mo agoHugging Face08s2e-lab /SecurityEval Dataset Card for SecurityEval This dataset is from the paper titled SecurityEval Dataset: Mining Vulnerability Examples to Evaluate Machine Learning-Based Code Generation Techniques. The project is accepted for The first edition of the International Workshop on Mining Software Repositories Applications for Privacy and Security (MSR4P&S '22). The paper describes the dataset for evaluating machine learning-based code generation output and the application of the dataset to the code… See the full description on the dataset page: https://huggingface.co/datasets/s2e-lab/SecurityEval.textn<1K10 likes305 downloads3y agoHugging Face09AYI-NEDJIMI /oauth-api-security-en OAuth & API Security Dataset (EN) Comprehensive English dataset covering OAuth 2.0 vulnerabilities, API attacks (OWASP API Top 10 2023), security controls, and Q&A pairs for training cybersecurity-specialized language models. Dataset Contents Category Entries Description OAuth 2.0 Vulnerabilities 20 Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage API Attacks 25 BOLA, BFLA, BOPLA, SSRF, GraphQL DoS, gRPC injection, CORS… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-en.textquestion-answeringn<1K0 likes246 downloads8mo agoHugging Face10natnitaract /exams-basic-and-quantum-cryptography-and-security-latex Open Problem Exams: Cryptography and Security (LaTeX) A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions. Dataset Overview Institution Files Topics Questions Caltech & TU Delft 8 38 145 EPFL 6 19 86 ETH Zurich 1 14 37 MIT 3 33 79 Total 18 104 347 Difficulty Distribution Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.textn<1K1 likes221 downloads7mo agoHugging Face11AYI-NEDJIMI /oauth-api-security-fr Dataset OAuth & Securite API (FR) Dataset francophone complet sur les vulnerabilites OAuth 2.0, les attaques API (OWASP API Top 10 2023), les controles de securite, et les questions-reponses pour l'entrainement de modeles de langage specialises en cybersecurite. Contenu du Dataset Categorie Nombre d'entrees Description Vulnerabilites OAuth 2.0 20 Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage Attaques API 25 BOLA, BFLA, BOPLA… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-fr.textquestion-answeringn<1K0 likes214 downloads8mo agoHugging Face12AndeXrd /SecurityQuestionstextn<1K0 likes171 downloads2y agoHugging Face13OpenClaw /clawhub-security-signals-live ClawHub Security Signals Live This dataset is the refreshed ClawHub security-signals corpus for scanner testing, prompt regression checks, and operational research against recent public ClawHub skills. It is a moving dataset, not the fixed paper benchmark. main is expected to change when the ClawHub security dataset snapshot workflow publishes a new sanitized export. Pin a Hugging Face revision or commit when you need reproducibility. For the frozen research-paper snapshot, use… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals-live.tabulartext-classification10K<n<100K0 likes161 downloads6d agoHugging Face14dattaraj /security-attacks-MITREtextquestion-answeringn<1K33 likes116 downloads2y agoHugging Face15adastracomputing /nixpkgs-security-patches nixpkgs-security-patches Training dataset for fine-tuning LLMs on nixpkgs security patch generation. Derived from real merged security PRs in NixOS/nixpkgs. Dataset Details 588 training examples / 66 eval examples (654 total) Format: Multi-turn tool-calling conversations in ChatML JSONL Each example is a realistic agent session: the model reads the package file, finds the upstream fix, computes hashes via tools, and submits the fix for approval Hashes and URLs… See the full description on the dataset page: https://huggingface.co/datasets/adastracomputing/nixpkgs-security-patches.texttext-generationn<1K1 likes114 downloads7mo agoHugging Face16sahilempire /redsec-security-sft-v1 RedSec Security SFT v1 A chat-formatted supervised fine-tuning dataset for security-focused language models, intended for authorized red-team, penetration-testing, and defensive use. Each record is a {"messages": [...]} conversation with an optional system turn, a user turn, and an assistant turn. Split Rows train 55,459 validation 1,155 test 1,155 total 57,769 Sources and attribution This is a derivative work. It combines, reformats… See the full description on the dataset page: https://huggingface.co/datasets/sahilempire/redsec-security-sft-v1.texttext-generation10K<n<100K0 likes112 downloads2mo agoHugging Face17TheFloatingString /s3_tf_s3_security_mcqtextn<1K0 likes106 downloads1y agoHugging Face18CaiZhiTech /Evaluation-Dataset-of-AI-Agent-Security-Guardrails DKnownAI Agent Security Evaluation Dataset Data Fields Field Type Description text string The adversarial input (prompt) to be evaluated by a security guardrail action string Human-annotated label: blocked or allowed Citation @misc{li2026comparativeevaluationaiagent, title={A Comparative Evaluation of AI Agent Security Guardrails}, author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.texttext-classification1K<n<10K1 likes100 downloads6mo agoHugging Face19logicBombExe /turkish_cyber_security_controls_benchmark Turkish Cyber Security Controls Benchmark Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir. v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5, Release 5.2.0 kontrol kataloğunu hedefler. Kapsam 100 Türkçe senaryo NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru 64 kontrol seçimi sorusu 17 denetim kanıtı sorusu 19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.textquestion-answeringn<1K4 likes100 downloads2mo agoHugging Face20davidfoss /bitcoin-security-reasoning-100k Dataset Card for Bitcoin Security Reasoning 100K 100,000 high-quality synthetic training samples for fine-tuning LLMs on Bitcoin protocol security analysis. Teaches models to analyze vulnerability clusters, form security hypotheses, and generate differential testing code. Dataset Details Dataset Description This dataset contains structured security reasoning chains for Bitcoin protocol vulnerabilities. Each sample presents a cluster of causal… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/bitcoin-security-reasoning-100k.texttext-generation100K<n<1M0 likes91 downloads9mo agoHugging Face21aicreatemo /clawhub-security-signals ClawHub Security Signals 🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale. This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree. Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/aicreatemo/clawhub-security-signals.tabulartext-classification10K<n<100K0 likes81 downloads21d agoHugging Face22MCPShield /mcp-security-scan-2026 MCP Security Scan Dataset 2026 Security scan results for 4,867 MCP (Model Context Protocol) server repositories, scanned by MCPShield. Dataset Description This is the largest public labeled MCP security dataset. Each entry contains the security grade, score, and detailed findings for a GitHub repository implementing an MCP server. Scanner MCPShield v5.0 — Two-pass detection architecture: Pass 1: 49 regex rules covering OWASP MCP Top 10 (94% detection on… See the full description on the dataset page: https://huggingface.co/datasets/MCPShield/mcp-security-scan-2026.tabulartext-classification1K<n<10K0 likes75 downloads6mo agoHugging Face23perchscan /security-pairs-public Perch Security Pairs — Public File: benchmark.jsonl — 1,311 JSONL records, one before/after method pair per record (2,622 method sides). Source code comes from open-source repositories; each repository's license applies to its code. Record fields Field Contents pair_id Stable pair identifier repository, language Source repository and code language source_datasets Source collection names before, after Method at the two source revisions… See the full description on the dataset page: https://huggingface.co/datasets/perchscan/security-pairs-public.texttext-classification1K<n<10K0 likes75 downloads2d agoHugging Face24bbkdevops /cass-vibe-security-bench CASS-Vibe Security Bench 19 teacher triples + 20 human-labeled rows + 100 real-CVE rows with reports. from datasets import load_dataset ds = load_dataset("bbkdevops/cass-vibe-security-bench", "truth20") Reproduce: python eval.py (needs the repo's cass/ package for scanners; numbers must match EVIDENCE.md scoreboard). texttext-classificationn<1K0 likes67 downloads19d agoHugging Face25burpsuite /Code_Vulnerability_Security_DPO Cybernative.ai Code Vulnerability and Security Dataset Dataset Description The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/burpsuite/Code_Vulnerability_Security_DPO.text1K<n<10K2 likes58 downloads9mo agoHugging Face26starknet-ai /cairo-security-audits Cairo Security Audits A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations. Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.tabulartext-retrievaln<1K1 likes56 downloads2mo agoHugging Face27rootly-ai-labs /terraform-s3-security-mcqtextn<1K0 likes53 downloads1y agoHugging Face28xanoutas /solidity-security-findings Solidity Security Findings Dataset 9,359 smart contract security findings from Code4rena audits. Statistics HIGH severity: 4,418 findings MEDIUM severity: 4,910 findings Sources: Code4rena, Immunefi Use Cases Training security audit models Vulnerability classification Smart contract analysis texttext-classification10K<n<100K3 likes53 downloads6mo agoHugging Face29dr3x1 /rizzo-pii-security-it rizzo-pii security IT — corpus sintetico del genere "sicurezza" Corpus sintetico italiano per la token classification di PII su documenti di sicurezza: verbali d'incidente, timeline forensi, ticket, estratti di log. Serve ad addestrare un modello che anonimizzi quei documenti in locale, prima di mandarli a un LLM esterno. È il dataset con cui è stato addestrato dr3x1/rizzo-pii-0.3B-security. Non contiene nessuna PII reale. Nessun dato personale vero, nessun indicatore di… See the full description on the dataset page: https://huggingface.co/datasets/dr3x1/rizzo-pii-security-it.texttoken-classification10K<n<100K0 likes53 downloads2mo agoHugging Face30referencesource /security-deposit-return-deadlines-by-state Security deposit return deadlines by US state — the statutory clock a landlord has to return a tenant's deposit, quoted from the state statute Canonical, always-current version: https://referencesource.org/security-deposit-return-deadlines-by-state/ Machine-readable: https://referencesource.org/security-deposit-return-deadlines-by-state/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-26 Stale after: 2027-08-26 (past this date, prefer the canonical copy —… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/security-deposit-return-deadlines-by-state.textn<1K0 likes50 downloads11d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.