Team Ai
Apppublic

pbhappliedsystems/quant-eval-agent-arena

sourceHugging Faceapache-2.0updated 1d agoView on Hugging Face
2likes
App README

quant-eval Agent Arena

Run two quantized models side by side on the same agent task, with each model's published behavioral results beside it.

Where the numbers come from

Every quanteval figure in this Space comes from `arenametrics.json`, written by the PBH Applied Systems site build from the published quant_eval Public Corpus (DOI 10.5281/zenodo.22009419) and, for showcase models, from sealed quant_eval run bundles published as summaries (per-case results on request). A model shows results only when the GGUF file this Space serves is the file a v7.22 run evaluated, matched by SHA-256 against the repo's own file metadata at startup.

Results pages

Models without a published v7.22 run show no scores here; their notes are labelled as v7.21 observations.

Community

Want a pair of models side by side that isn't here yet, or have results of your own to compare?

<!-- pbh-community -->

Request an evaluation · Live agent demo