Team Ai
Datasetpublicgated

aporthq/vault-benchmark-v1

APort Vault Benchmark v1 4,371 attacks written by humans in 1,128 sessions of a public competition (March to August 2026), replayed against 14 language models from 8 labs acting as a simulated bank teller (no real money moves), in two conditions: model alone, and behind a pre-action authorization layer. Papers APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport (Paper 2). The replay benchmark and results reported in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/aporthq/vault-benchmark-v1.

sourceHugging Facecc-by-4.0updated 20d agoView on Hugging Face
2likes171downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.