Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K18 likes17k downloads4mo agoHugging Face02Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads3d agoHugging Face03typesafe /evalsafe-security-incidents Security incidents Snapshot: 2026-09-28. 240 cases and 1,820 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-security-incidents.tabular1K<n<10K2 likes1.1k downloads11d agoHugging Face04secmlr /llm-fv-security-targets LLM-FV Security Targets This dataset contains 889 independently validated, containerized security-agent targets produced by the ucsb-mlsec/llm-fv pipelines. Contents GitHub Global Security Advisories: 446 targets OSS-Fuzz: 442 targets PoC task support: 889 targets Exploit task support: 447 targets Patch task support: 447 targets Compressed bundle size: 42.41 GiB Vulnerability classes: {'logic_bug': 450, 'memory_vulnerability': 439} Primary languages: {'C': 118… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/llm-fv-security-targets.tabulartext-generationn<1K0 likes1.1k downloads2d agoHugging Face05ZipLime /security-master US Security Master Which company a ticker belonged to, on a date. Which filer a CUSIP points at. What a company used to be called. 23 553 filers · 144 054 identifier spans · 10 061 784 dated observations · 74 814 CUSIP–ticker pairs · 1960 to 2026 This is the glue for the rest of the family. Every other ZipLime dataset keys on the SEC's CIK, which never changes — and every price series, broker feed and research note keys on a ticker, which changes all the time. Joining the two… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/security-master.tabulartabular-classification10M<n<100M0 likes587 downloads3d agoHugging Face06Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes512 downloads3mo agoHugging Face07Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M1 likes440 downloads1mo agoHugging Face08Publicus /cvefixes-security-ir-graphrag CVEfixes Security IR GraphRAG This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup. All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.tabular100K<n<1M0 likes438 downloads2mo agoHugging Face09OpenClaw /clawhub-security-signals ClawHub Security Signals 🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale. This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree. Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals.tabulartext-classification10K<n<100K53 likes343 downloads4mo agoHugging Face10stacklok /llm-security-leaderboard-contentstabularn<1K0 likes278 downloads1y agoHugging Face11CyberMax-tools /saas-email-security-snapshot DomainDNA: SaaS email security & tech stack snapshot Check SPF, DMARC and MX for your own domains, by API: $5 for 1,000 checks. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com Need it fresh, filtered or via API? This free file is a snapshot (one scan of SaaS domains' email security), last updated 2026-09-23. Mailvett email and domain check API (50 free calls a day, $5 for… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/saas-email-security-snapshot.tabulartabular-classificationn<1K0 likes244 downloads2d agoHugging Face12Cyber-security-final-project /HARMLESS_Synthetic_Injected_PDFs_EDA Injected PDFs - EDA and Evaluation Corpus This repository holds the exploratory data analysis for a project on detecting harmless-but-real attack payloads injected into PDF files, together with the dataset that analysis produced. The project has two halves, both in the notebook Final_project_V7_EDA.ipynb: Question Input Part 1 Is our synthetic corpus a stand-in for real malware, or is it something else? The published CIC feature table (11,126 x 34) Part 2 Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.imagetext-classification1K<n<10K0 likes166 downloads2mo agoHugging Face13OpenClaw /clawhub-security-signals-live ClawHub Security Signals Live This dataset is the refreshed ClawHub security-signals corpus for scanner testing, prompt regression checks, and operational research against recent public ClawHub skills. It is a moving dataset, not the fixed paper benchmark. main is expected to change when the ClawHub security dataset snapshot workflow publishes a new sanitized export. Pin a Hugging Face revision or commit when you need reproducibility. For the frozen research-paper snapshot, use… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals-live.tabulartext-classification10K<n<100K0 likes161 downloads5d agoHugging Face14Cyber-security-final-project /Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition Injected PDFs - Model Evaluation This repository holds the model evaluation stage of a project on detecting harmless-but-real attack payloads injected into PDF files, together with the artefacts it produced for the application. Nothing is trained here. Seven off-the-shelf models are measured against the same 1,100 PDFs, and the two winners are exported for the app to load. Question Candidates Winner Part A Which files look like this one? 3 embedding models x 2 inputs… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition.tabulartext-classification1K<n<10K0 likes117 downloads2mo agoHugging Face15rgaucher /security-vuln-patches Security Vulnerability Fix Patches A curated dataset of 355 validated real-world vulnerability fix patches mined from public GitHub repositories. Each record contains the unified diff of a pull request or commit that fixes a specific CWE (Common Weakness Enumeration) vulnerability, along with metadata about the repository, language, and fix type. Dataset Summary Source: Public GitHub PRs and commits (2024-2025) discovered via BigQuery analysis of GH Archive data… See the full description on the dataset page: https://huggingface.co/datasets/rgaucher/security-vuln-patches.tabulartext-classificationn<1K0 likes99 downloads6mo agoHugging Face16BAAI /IndustryCorpus2_other_information_services_information_security IndustryCorpus2: Information Services This repository contains the IndustryCorpus2: Information Services domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}, year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_other_information_services_information_security.tabular100K<n<1M1 likes92 downloads2mo agoHugging Face17jasonzhuyansen /agent-skills-security-grades Agent Skills Security Grades Security grades and quality scores for 130,173 open-source AI agent skills and MCP servers collected from GitHub, from Agent Skills Hub. Each row is one skill/server with a rule-based security grade (SAFE / CAUTION / UNSAFE / REJECT / UNAUDITED), red-flag identifiers, and a 0–100 quality score. Why this exists AI coding agents install third-party skills that run with the agent's full permissions and credentials, but marketplaces rank… See the full description on the dataset page: https://huggingface.co/datasets/jasonzhuyansen/agent-skills-security-grades.tabulartabular-classification100K<n<1M1 likes89 downloads3mo agoHugging Face18electricsheepafrica /africa-synth-agriculture-and-food-security-dataset-all African Agriculture & Food Security Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory) Size category: 1M<n<10M - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-agriculture-and-food-security-dataset-all.tabulartabular-classification1M<n<10M0 likes87 downloads2mo agoHugging Face19ScortonAI /Cyber-Security-Breachestabular1K<n<10K12 likes85 downloads3y agoHugging Face20Publicus /cvefixes-security-autoformal-span-cache CVEfixes Security source spans This dataset contains exact, source-bound prose, code, and diff spans derived from the original-data configuration of Publicus/cvefixes-security-ir-graphrag at revision 6fd5918bed34f8851430e74a149502587a953fe2. The underlying source is hitoshura25/cvefixes at revision d4f5c4ea65329d9ccbb8a3b3149e5d06eda5edb2. The extraction considered all 12,987 original rows across all three shards. It retained 9,402 source rows and excluded 3,585. The release… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-autoformal-span-cache.tabulartext-generation100K<n<1M0 likes83 downloads10d agoHugging Face21aicreatemo /clawhub-security-signals ClawHub Security Signals 🦀 ClawHub | 📝 OpenClaw Blog | 🤗 Hugging Face Blog | 📄 Paper | 📄 Pre-Print ClawHub Security Signals is a sanitized, MIT-licensed security-signals dataset for public OpenClaw agent skills. It captures how an agent-skill registry evaluates trust, provenance, bundled code, and scanner evidence at scale. This dataset was presented in the paper ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree. Paper snapshot: this… See the full description on the dataset page: https://huggingface.co/datasets/aicreatemo/clawhub-security-signals.tabulartext-classification10K<n<100K0 likes81 downloads21d agoHugging Face22Nobody05 /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/cyber-security-100m.tabulartext-classification100M<n<1B0 likes80 downloads3mo agoHugging Face23MCPShield /mcp-security-scan-2026 MCP Security Scan Dataset 2026 Security scan results for 4,867 MCP (Model Context Protocol) server repositories, scanned by MCPShield. Dataset Description This is the largest public labeled MCP security dataset. Each entry contains the security grade, score, and detailed findings for a GitHub repository implementing an MCP server. Scanner MCPShield v5.0 — Two-pass detection architecture: Pass 1: 49 regex rules covering OWASP MCP Top 10 (94% detection on… See the full description on the dataset page: https://huggingface.co/datasets/MCPShield/mcp-security-scan-2026.tabulartext-classification1K<n<10K0 likes75 downloads6mo agoHugging Face24HashgraphOnline /hol-plugin-security HOL Plugin Security Snapshot of Hashgraph Online plugin-catalog scores, modeled HOL Guard runtime fixtures, and curated public advisories. Dataset id: HashgraphOnline/hol-plugin-security. The default config is plugins (205 scored registry plugins). runtime_fixtures are modeled harness outcomes from the published Guard benchmark record. advisories are the 22 public HOL Guard advisory pages. advisories.threat_class is set from the official hub listing badges on… See the full description on the dataset page: https://huggingface.co/datasets/HashgraphOnline/hol-plugin-security.tabularn<1K1 likes60 downloads27d agoHugging Face25starknet-ai /cairo-security-audits Cairo Security Audits A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations. Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.tabulartext-retrievaln<1K1 likes56 downloads2mo agoHugging Face26Dhanjo /ai-agent-security-dataset AI Agent Security and System Prompt Leakage Dataset Dataset Overview This dataset was created for research on AI agent security, with a specific focus on system prompt leakage, jailbreak resistance, and security-aligned fine-tuning. The dataset evaluates how often AI agents reveal confidential information embedded inside their system prompts when exposed to adversarial prompts. It also compares the behavior of a baseline language model against a model fine-tuned using… See the full description on the dataset page: https://huggingface.co/datasets/Dhanjo/ai-agent-security-dataset.tabulartext-generation1K<n<10K1 likes53 downloads5mo agoHugging Face27rogue-security /real-world-benign-use-cases Real-World Benign Use Cases A curated set of 178 real-world, 100%-benign examples (label == 0 for every row) pulled from production AI-coding-agent traffic — chat messages, tool output, shell commands, code snippets — built specifically to stress-test prompt-injection / jailbreak classifiers for false positives. Every row was independently judged benign with high confidence before inclusion. This is not a random sample of production traffic: rows were preferentially drawn from… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/real-world-benign-use-cases.tabulartext-classificationn<1K0 likes51 downloads2mo agoHugging Face28Neura-parse /quantum-cryptography-and-post-quantum-security Neura Parse — Quantum Cryptography & Post-Quantum Security A deep vertical on cryptography that uses quantum mechanics and on classical cryptography built to resist quantum attack. It covers quantum key distribution (BB84, B92, six-state, SARG04, E91, BBM92, decoy-state, MDI-QKD, TF-QKD, CV-QKD), device-independent protocols, composable and finite-key security proofs, quantum hacking with countermeasures, classical post-processing (reconciliation, privacy amplification… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-cryptography-and-post-quantum-security.tabularquestion-answering100K<n<1M0 likes50 downloads3mo agoHugging Face29electricsheepafrica /africa-suite-of-food-security-indicators-value Suite of Food Security Indicators — Value | Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-suite-of-food-security-indicators-value.tabulartabular-classification10K<n<100K0 likes46 downloads2mo agoHugging Face30ammarnasr /Python-Security-Code-Datasettabular1K<n<10K3 likes45 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.