Team Ai
Datasetpublic

pravinai/agent-reliability-eval

Agent Reliability Eval A short, runnable notebook on evaluating agent reliability along two axes scored separately: tool-call accuracy and hallucination rate (groundedness of the final answer against what the agent's tools actually returned). Open agent_reliability_eval.ipynb — it runs end to end with no API key and no external dependencies beyond nbformat/nbclient if you want to re-execute it; the shipped copy already has outputs baked in. Why two metrics instead of… See the full description on the dataset page: https://huggingface.co/datasets/pravinai/agent-reliability-eval.

sourceHugging Faceupdated 28d agoView on Hugging Face
0likes32downloads
Dataset Card

Agent Reliability Eval

A short, runnable notebook on evaluating agent reliability along two axes scored separately: tool-call accuracy and hallucination rate (groundedness of the final answer against what the agent's tools actually returned).

Open agent_reliability_eval.ipynb — it runs end to end with no API key and no external dependencies beyond nbformat/nbclient if you want to re-execute it; the shipped copy already has outputs baked in.

Why two metrics instead of one

End-to-end accuracy hides two different failure modes behind a single number: an agent that calls the wrong tool but gets lucky, and an agent that calls the right tool but pads its answer with unsupported specifics. This notebook scores three mock agent policies with the same harness to show it actually discriminates between those failure profiles, not just pass/fail.

It also documents, on purpose, a real blind spot in the grounding check used here — a fabricated qualitative claim the simple heuristic misses — because pointing at your own eval's limitations is more useful than hiding them.

Companion project

`auditable-agent-demo` applies the same "don't trust output without checking how it got there" principle to a spec-bound, auditable compliance-triage agent with a full tool-call and rule-citation trail — the production-shaped version of the idea this notebook explores in miniature.

Files

  • —agent_reliability_eval.ipynb — the notebook (run it, or read the baked-in outputs).
  • —reliability_eval.py — the underlying logic (tools, mock policies, eval functions), importable on its own.