Leanmcp/lawfulbench
LAWFUL-Bench LAWFUL-Bench: Measuring Whether LLM Agents Apply Data Protection Law Dheeraj Pai, Lu Xian (Leanmcp) An agentic benchmark for operational data protection duties under the GDPR. An agent under test and a simulated data subject each hold tools over one shared database, and 44 documents of primary law are reachable through retrieval rather than pasted into the prompt. The graded artifact is a justification triple -- (decision, lawful_basis, record_action) -- filed… See the full description on the dataset page: https://huggingface.co/datasets/Leanmcp/lawfulbench.
LAWFUL-Bench
LAWFUL-Bench: Measuring Whether LLM Agents Apply Data Protection Law Dheeraj Pai, Lu Xian (Leanmcp)
An agentic benchmark for operational data protection duties under the GDPR. An agent under test and a simulated data subject each hold tools over one shared database, and 44 documents of primary law are reachable through retrieval rather than pasted into the prompt.
The graded artifact is a justification triple -- (decision, lawful_basis, record_action) -- filed through a compliance tool, alongside the resulting world state and what the person was actually told. An agent that reaches the right outcome for the wrong reason is not compliant, and this benchmark scores that difference.
- Dataset: <https://huggingface.co/datasets/Leanmcp/lawfulbench>
- Paper: Under review at ARR/EACL. Link added here when the preprint is posted.
- Code / harness: released separately, link added here when public.
Read this before using it
- Gold labels have not been reviewed by a qualified data protection practitioner.
- No human baseline exists, so absolute scores are uninterpretable.
- Redistribution terms for the EDPB/WP29 document are unconfirmed (excluded for now).
- No inter-annotator agreement has been measured on any item.
Treat these numbers as a research artifact, not as a compliance assurance. A high score is evidence about a model on 24 constructed items. It is not evidence that a deployment is lawful, and it must not be cited as such.
Composition
24 tasks across four stateful domains (bank_support, hospital, hr_enterprise, personal_assistant), 6 each.
Difficulty: easy 1, edge 5, hard 9, medium 9.
Files
Quick start
from datasets import load_dataset
ds = load_dataset("Leanmcp/lawfulbench", split="train")
print(ds[0]["id"], ds[0]["class"], ds[0]["domain"])The world state, the policy text and the retrieval corpus are plain files rather than dataset columns, so fetch them alongside:
from huggingface_hub import snapshot_download
path = snapshot_download("Leanmcp/lawfulbench", repo_type="dataset")
# path/domains/<d>/db.json, path/domains/<d>/policy.txt, path/docs/, path/MANIFEST.jsonA task is only reproducible against the db.json it shipped with: the gold triple is defined relative to that initial world.
Metrics
The column names in results/summary.csv, in the order they matter.
Baseline
1 agents, one trial each, user simulator and judge held constant.
Justified accuracy stays near the floor while decision accuracy runs two to four times higher: most correct decisions are reached on a wrong legal basis. Over-refusal exceeds under-refusal in every model tested.
Licensing
- GDPR and EU AI Act text: EUR-Lex, reuse permitted with attribution (Commission Decision 2011/833/EU).
- Sector retention rules (
docs/sector/): synthetic, written by the authors, not attributable to any real institution, and not authoritative law. - Tasks, databases, policies: released by the authors under the repository's licence.
- The EDPB/WP29 guidance document used in the harness is excluded from this release pending confirmation of redistribution terms.
MANIFEST.jsoncarries its source URL and hash so it can be fetched locally.
Personal data
None. Every data subject, record and identifier in the four databases is fictitious. No real personal data was collected, processed, or stored at any point in constructing this benchmark.
Running it
The harness is not in this dataset repository. See the code release for lawful run / lawful eval, the four tool surfaces, and the judge prompts.
Citation
If you use LAWFUL-Bench, please cite the paper:
@misc{pai2026lawfulbench,
title = {LAWFUL-Bench: Measuring Whether LLM Agents Apply Data Protection Law},
author = {Dheeraj Pai and Lu Xian},
year = {2026},
note = {Under review at ARR/EACL},
howpublished = {Hugging Face dataset, \url{https://huggingface.co/datasets/Leanmcp/lawfulbench}}
}Contact
Dheeraj Pai and Lu Xian, Leanmcp. Corrections to a gold label are welcome and wanted: open a discussion on this repository with the task id and the Article you think it turns on.
Generated from commit eac7cd1.
