Team Ai
Datasetpublic

libredb/database-agent-runs

LibreDB Agent Benchmark 8,199 agent runs · 39 open-weight models served locally, plus one hosted model as a control · 6 task surfaces · 110,711 ledger events · 14,008 refused tool calls This is the complete measurement record behind the paper What Stops a Small Language Model From Driving a Database Agent. It is not a scored summary: it is every event the server wrote while the runs happened, released so that every number in the paper can be recomputed, and disagreed with… See the full description on the dataset page: https://huggingface.co/datasets/libredb/database-agent-runs.

sourceHugging Facemitupdated 19d agoView on Hugging Face
4likes439downloads
7 commits on main
b3659a719d ago

Tag the dataset with the arXiv paper it accompanies

cevheri
c4f17e919d ago

Record the announced arXiv id, correct two stale refusal counts, name the taxonomy population

cevheri
7ba04ee22d ago

Upload verify.py with huggingface_hub

cevheri
a617a8b23d ago

Verifier: add the agent-mode taxonomy, the planning-is-toolless checks and the intervention table

cevheri
2f5ed2223d ago

Card: cite the dataset DOI, add the paper citation, correct the exporter's location

cevheri
54903c523d ago

Add the measurement corpus: 8,199 runs, 110,711 ledger events, 14,008 refusals, with scorer and verifier

cevheri
e4f29de23d ago

initial commit

cevheri