libredb/database-agent-runs
LibreDB Agent Benchmark 8,199 agent runs · 39 open-weight models served locally, plus one hosted model as a control · 6 task surfaces · 110,711 ledger events · 14,008 refused tool calls This is the complete measurement record behind the paper What Stops a Small Language Model From Driving a Database Agent. It is not a scored summary: it is every event the server wrote while the runs happened, released so that every number in the paper can be recomputed, and disagreed with… See the full description on the dataset page: https://huggingface.co/datasets/libredb/database-agent-runs.
4439
