agent-security
kio12-security-agent-eurollm-9b-instruct-2512-v2kio12-security-agent-ministral3-8b-reasoning-2512-v2kio12-security-agent-gemma4-12b-v2kio12-security-agent-foundation-sec-8b-reasoning-v2kio12-security-agent-qwen35-9b-v2kio12-security-agent-qwen25-coder-7b-v2gemma-2-2b-agent-securityQwen-Security-Agent-3B
White-Hat-Security-Agent-Prompts-600K
White Hat Security Agent Prompts 600K
Overview
The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios.
Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.agent-skill-security-paper-artifacts
Agent Skill Security Research Artifacts
This collection releases derived aggregate evidence, original figure data and plots, offline analysis code, and reproducibility manifests. The current research portfolio has 4 consolidated empirical manuscript directions: scanner configuration/gate comparisons, small-model score interfaces and source transfer, runtime cascade action composition, and deterministic-rule maintenance/evidence contracts. Current authoring/delivery state: Four… See the full description on the dataset page: https://huggingface.co/datasets/Vineethsain/agent-skill-security-paper-artifacts.coding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.autonomous-ai-agent-security-incidents-2026
Autonomous AI Agent Security Incidents of 2026: Benchmark Dataset & Incident Corpus
An open, verifiable, machine-readable dataset documenting 109 empirical security and containment failures involving autonomous AI coding, orchestration, and execution agents observed during 2026.
This dataset accompanies the 689-page open-access research monograph published on Zenodo:
Doletskyi, Serhii. (2026). Autonomous AI Agent Security Incidents of 2026: Empirical Incident Corpus, Sandbox… See the full description on the dataset page: https://huggingface.co/datasets/doletskyisergey/autonomous-ai-agent-security-incidents-2026.agent-security-datasets
prompt-protection datasets
A held-out, human-authored corpus for evaluating prompt-injection / agent-security
guards. Every item is original to this project and disjoint from the unit-test
fixtures (tests/__fixtures__/*.txt) and bench corpus (bench/corpus/*), so it measures
generalisation rather than memorisation. JSONL, one object per line, UTF-8.
Files
agent-flows.jsonl, tool-call guard scenarios (100: 50 attack / 50 benign)
Each row is a full… See the full description on the dataset page: https://huggingface.co/datasets/promptprotection/agent-security-datasets.
