Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K17 likes16k downloads4mo agoHugging Face02gussieIsASuccessfulWarlock /security_instruct_mcq_2481textn<1K0 likes4.5k downloads2y agoHugging Face03Vyber07 /cyber-securitygated Cybersecurity AI Knowledge Base — PhD-Level Dataset Overview This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security. Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms Purpose Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.texttext-generationn<1K116 likes3.4k downloads1mo agoHugging Face04kala185 /comptia_security_pluse_701text1K<n<10K0 likes2.7k downloads1y agoHugging Face05Nikolife /pulsefeed-x402-security PulseFeed — x402 Agent-Payment Security & Trust (open data) Independent, daily-updated trust & safety data for the x402 agent-payment economy (HTTP 402 + stablecoins on Base) and the MCP server ecosystem — by PulseFeed. AI agents increasingly pay for APIs autonomously over x402 and connect to MCP servers that can run code on install. But 18% of listed x402 endpoints are dead or invalid, and "live" is not the same claim as "payable": of 32095 endpoints that return a valid 402… See the full description on the dataset page: https://huggingface.co/datasets/Nikolife/pulsefeed-x402-security.10K<n<100K0 likes2.4k downloads1h agoHugging Face06yatin-superintelligence /White-Hat-Security-Agent-Prompts-600K White Hat Security Agent Prompts 600K Overview The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios. Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.texttext-generation100K<n<1M21 likes2.1k downloads7mo agoHugging Face07pkgforge-security /domains Internet Domains Domains HuggingFace Hub Mirror for https://github.com/pkgforge-security/domains The Sync Workflow actions are at: https://github.com/pkgforge-security/domains TOS & Abuse (To Hugging-Face's Staff) Hi, if you are an offical from Hugging-Face here to investigate why this Repo is so Large and are considering deleting, & terminating our Account. Please note that, this project benefits a lot of people (You can do a… See the full description on the dataset page: https://huggingface.co/datasets/pkgforge-security/domains.text10B<n<100B5 likes1.8k downloads2mo agoHugging Face08James4Ever0 /network_security_questionsThis dataset contains a single file full of network security questions in Chinese. Could be used as good initial sources for scrapers, though not good as your browsing history. text1M<n<10M10 likes1.4k downloads3y agoHugging Face09rogue-security /prompt-injections-benchmarkgated Dataset: Qualifire Benchmark Prompt Injection(Jailbreak vs. Benign) Datasets Overview This dataset contains 5,000 prompts, each labeled as either jailbreak or benign. The dataset is designed for evaluating AI models' robustness against adversarial prompts and their ability to distinguish between safe and unsafe inputs. Dataset Structure Total Samples: 5,000 Labels: jailbreak, benign Columns: text: The input text label: The classification (jailbreak or benign)… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/prompt-injections-benchmark.text1K<n<10K45 likes1.2k downloads6mo agoHugging Face10CyberNative /Code_Vulnerability_Security_DPO Cybernative.ai Code Vulnerability and Security Dataset Dataset Description The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen on… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative/Code_Vulnerability_Security_DPO.text1K<n<10K172 likes1.1k downloads3y agoHugging Face11pAILabs /infosec-security-qatext10K<n<100K12 likes1.1k downloads2y agoHugging Face12Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K7 likes984 downloads11mo agoHugging Face13oi-uae /cyber-securitygated Cybersecurity Instruction-Tuning Dataset A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning, built from 198 distinct sources spanning offensive security, blue-team operations, vulnerability intelligence, cloud/AWS security, malware analysis, digital forensics, and more. Every record is normalized to the standard messages chat format and deduplicated at both file and record level. ⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.textquestion-answering1M<n<10M27 likes945 downloads24d agoHugging Face14OliseNS /person-face-package-home-security-detection Person, face & package — home-security detection dataset YOLO-format dataset for person, face, vehicles, small vehicles, parcels, pets, and birds in home / delivery / street scenes. Full data lives under bigsplit/; sample/ is a 100-image preview subset (same layout: flat images/ and labels/). Training configs point at a data.yaml beside those folders. Both splits use the same class list (nc and names); only the root path and which image set is packaged differ. Labels are YOLO .txt… See the full description on the dataset page: https://huggingface.co/datasets/OliseNS/person-face-package-home-security-detection.imageobject-detectionn<1K1 likes913 downloads6mo agoHugging Face15domblake /airport-securityimagen<1K0 likes860 downloads3y agoHugging Face16taroii /airport-security-detectionimagen<1K0 likes842 downloads3y agoHugging Face1711-47 /Organized_PreTrain_Cyber_Security_640ktext100K<n<1M0 likes756 downloads2mo agoHugging Face18Publicus /cvefixes-security-ir-graphrag CVEfixes Security IR GraphRAG This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup. All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.tabular100K<n<1M0 likes698 downloads2mo agoHugging Face19typesafe /evalsafe-security-incidents Security incidents Snapshot: 2026-09-28. 240 cases and 1,820 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-security-incidents.tabular1K<n<10K2 likes689 downloads7d agoHugging Face20secmlr /llm-fv-security-targets LLM-FV Security Targets This dataset contains 1183 independently validated, containerized security-agent targets produced by the ucsb-mlsec/llm-fv pipelines. Contents GitHub Global Security Advisories: 394 targets OSS-Fuzz: 788 targets PoC task support: 1183 targets Exploit task support: 395 targets Patch task support: 395 targets Compressed bundle size: 58.51 GiB Vulnerability classes: {'logic_bug': 409, 'memory_vulnerability': 774} Primary languages: {'C': 201… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/llm-fv-security-targets.tabulartext-generation1K<n<10K0 likes639 downloads7h agoHugging Face21s0u9ata /security-kg Security Knowledge Graph Triples Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis. Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/s0u9ata/security-kg.textgraph-ml10M<n<100M0 likes617 downloads14h agoHugging Face22Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M1 likes563 downloads1mo agoHugging Face23ZipLime /security-master US Security Master Which company a ticker belonged to, on a date. Which filer a CUSIP points at. What a company used to be called. 23 553 filers · 144 054 identifier spans · 10 061 784 dated observations · 74 814 CUSIP–ticker pairs · 1960 to 2026 This is the glue for the rest of the family. Every other ZipLime dataset keys on the SEC's CIK, which never changes — and every price series, broker feed and research note keys on a ticker, which changes all the time. Joining the two… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/security-master.tabulartabular-classification10M<n<100M0 likes517 downloads20h agoHugging Face24NamkinBhujiya /network_security_questionsThis dataset contains a single file full of network security questions in Chinese. Could be used as good initial sources for scrapers, though not good as your browsing history. text1M<n<10M0 likes510 downloads4mo agoHugging Face25heron-ai-security /stegoattack-advbench50 StegoAttack AdvBench-50 Steganographic jailbreak data generated using the StegoAttack pipeline from the paper "Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks" (Geng et al., 2025). For experiment results and analysis, see experiment.md. What is StegoAttack? StegoAttack is a jailbreak method that uses steganography to hide harmful queries inside benign-looking text. It embeds each word of a harmful query at a fixed position (e.g. the 2nd… See the full description on the dataset page: https://huggingface.co/datasets/heron-ai-security/stegoattack-advbench50.text-generationn<1K0 likes465 downloads20d agoHugging Face26Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes387 downloads3mo agoHugging Face27ayshajavd /code-security-vulnerability-dataset Code Security Vulnerability Dataset A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories. Dataset Details Property Value Total Samples 175,419 Train / Val / Test 140,335 / 17,542 / 17,542 Languages C, C++, Python, JavaScript, Java, PHP, Go Labels 31 (multi-label) Format Parquet with… See the full description on the dataset page: https://huggingface.co/datasets/ayshajavd/code-security-vulnerability-dataset.texttext-classification100K<n<1M7 likes380 downloads6mo agoHugging Face28s2e-lab /SecurityEval Dataset Card for SecurityEval This dataset is from the paper titled SecurityEval Dataset: Mining Vulnerability Examples to Evaluate Machine Learning-Based Code Generation Techniques. The project is accepted for The first edition of the International Workshop on Mining Software Repositories Applications for Privacy and Security (MSR4P&S '22). The paper describes the dataset for evaluating machine learning-based code generation output and the application of the dataset to the code… See the full description on the dataset page: https://huggingface.co/datasets/s2e-lab/SecurityEval.textn<1K10 likes369 downloads3y agoHugging Face29AndeXrd /SecurityQuestionstextn<1K0 likes355 downloads2y agoHugging Face30Cyber-security-final-project /Generated_Injected_PDFs_HARMLESS Generated Injected PDFs — HARMLESS A synthetic dataset of 1,100 PDF files built for training and evaluating structural PDF-malware detectors. It pairs benign PDFs with PDFs into which safe, non-executable "malware-shaped" objects have been injected, so a model can learn to separate the two from byte-level structure alone. ⚠️ Safety notice — read first Nothing in this dataset is real malware. Every injected payload is built from industry-standard, non-executable… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Generated_Injected_PDFs_HARMLESS.documenttabular-classification1K<n<10K0 likes349 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.