aradhye/agent-safety-bench
Agent Safety Bench (ASB) ASB is a benchmark for evaluating the safety of tool-using LLM agents. Each example pairs a natural-language instruction with one or more sandboxed tool environments; the goal is to measure whether an agent completes the task without taking unsafe actions. This repository hosts the task data for ASB. The runtime environments themselves (the Python classes the agent calls into) live in the companion package agent-safety-bench-envs. It ships two configs:… See the full description on the dataset page: https://huggingface.co/datasets/aradhye/agent-safety-bench.
Agent Safety Bench (ASB)
ASB is a benchmark for evaluating the safety of tool-using LLM agents. Each example pairs a natural-language instruction with one or more sandboxed tool environments; the goal is to measure whether an agent completes the task without taking unsafe actions.
This repository hosts the task data for ASB. The runtime environments themselves (the Python classes the agent calls into) live in the companion package `agent-safety-bench-envs`.
It ships two configs:
default— the core ASB tasks (3,464 rows). One row per task, paired with the env specs the agent will operate on.augmented— synthetically generated training tasks (756 rows). Each row carriespython_codesfor fresh env classes alongside the matchingtool_descsandparameters; useful as a training-time mixin.
Both configs are exposed under the train split. ASB has no held-out evaluation split here — downstream consumers do their own split if needed.
Quickstart
import json
from datasets import load_dataset
from asb_envs import EnvManager
# Default config — core ASB tasks
ds = load_dataset("aradhye/agent-safety-bench", split="train")
mgr = EnvManager()
example = ds[0]
envs = json.loads(example["environments"])
env_spec = envs[0]
env = mgr.init_env(env_spec["name"], env_spec.get("parameters") or None)
print(example["instruction"])
print("tools:", env.tool_list)
result = env.call_tool("send_email", {"receiver": ["a@b.com"], "content": "hi"})
print(result)
# Augmented config — synthetic training tasks
aug = load_dataset("aradhye/agent-safety-bench", "augmented", split="train")
print(aug[0]["instruction"])Schema
default config (split train, 3,464 rows)
environments and dialog are JSON-encoded because the inner shapes (in particular parameters) vary per environment, which doesn't fit a single Arrow schema cleanly. Decode with json.loads() after loading.
augmented config (split train, 756 rows)
Each row may define multiple environments; the three list fields are position-aligned. Synthetic envs share names with default envs only by coincidence — treat them as distinct from the agent-safety-bench-envs package and instantiate from python_codes if you need them at runtime.
License
Apache License 2.0. See LICENSE in the source repo.
