datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
basilisk-webpentest
Basilisk WebPentest Dataset
Instruction‑tuning data for web penetration testing, used to train sl4de/Basilisk-7B.
Direct, uncensored expert answers (payloads, PoCs, tool commands, reports) for authorized security testing and education.
Files
File
Records
Domains
basilisk_v0.jsonl
2,999
SQLi, XSS, Recon
basilisk_v1.jsonl
8,054
19 (v0 + SSRF, XXE, deserialization, auth/JWT, access control, API, SSTI, LFI, CSRF, CORS, request smuggling, prototype… See the full description on the dataset page: https://huggingface.co/datasets/sl4de/basilisk-webpentest.WEBPRMBENCH
WebPRMBench
The first comprehensive evaluation benchmark for Web Process Reward Models
Published at ICLR 2026
Paper | Code | Website | Collection | Demo
Overview
WebPRMBench is the first comprehensive evaluation benchmark dedicated to Web Process Reward Models (WebPRMs). It evaluates how well a reward model can judge the quality of web agent actions during long-horizon web navigation. Each instance presents a web state (page context, trajectory history, user… See the full description on the dataset page: https://huggingface.co/datasets/ZYao720/WEBPRMBENCH.
