Team Ai
Datasetpublic

ruchit11111/coding-agent-security-benchmark

Coding Agent Security Benchmark A benchmark for evaluating whether an LLM can correctly identify security violations in the behavior of an autonomous coding agent - spanning dangerous shell commands, credential leakage, prompt injection, supply-chain risk, privacy leaks, and more. Each row is a single message sampled from a coding-agent session (a user instruction, a tool call the agent issued, a tool's response, or the agent's own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/ruchit11111/coding-agent-security-benchmark.

sourceHugging Facecc-by-nc-4.0updated 1mo agoView on Hugging Face
1likes74downloads
1 commits on main
a3bc5281mo ago

Duplicate from rogue-security/coding-agent-security-benchmark

ruchit11111, dror44