libredb/database-agent-runs
LibreDB Agent Benchmark 8,199 agent runs · 39 open-weight models served locally, plus one hosted model as a control · 6 task surfaces · 110,711 ledger events · 14,008 refused tool calls This is the complete measurement record behind the paper What Stops a Small Language Model From Driving a Database Agent. It is not a scored summary: it is every event the server wrote while the runs happened, released so that every number in the paper can be recomputed, and disagreed with… See the full description on the dataset page: https://huggingface.co/datasets/libredb/database-agent-runs.
Tag the dataset with the arXiv paper it accompanies
Record the announced arXiv id, correct two stale refusal counts, name the taxonomy population
Upload verify.py with huggingface_hub
Verifier: add the agent-mode taxonomy, the planning-is-toolless checks and the intervention table
Card: cite the dataset DOI, add the paper citation, correct the exporter's location
Add the measurement corpus: 8,199 runs, 110,711 ledger events, 14,008 refusals, with scorer and verifier
initial commit
