pbhappliedsystems/quant-eval-agent-arena
quant-eval Agent Arena
Run two quantized models side by side on the same agent task, with each model's published behavioral results beside it.
Where the numbers come from
Every quanteval figure in this Space comes from `arenametrics.json`, written by the PBH Applied Systems site build from the published quant_eval Public Corpus (DOI 10.5281/zenodo.22009419) and, for showcase models, from sealed quant_eval run bundles published as summaries (per-case results on request). A model shows results only when the GGUF file this Space serves is the file a v7.22 run evaluated, matched by SHA-256 against the repo's own file metadata at startup.
Results pages
- Ministral-3-14B-Instruct-2512 Q4KM: pbhappliedsystems.com/models/ministral-3-14b-instruct-2512#Q4_K_M — shown once the repo serves the evaluated file
- Mistral-Nemo-Instruct-2407 Q4KM: pbhappliedsystems.com/models/mistral-nemo-instruct-2407#Q4_K_M
- Qwen2.5-14B-Instruct-1M Q4KM: pbhappliedsystems.com/models/qwen2.5-14b-instruct-1m#Q4_K_M
- Qwen2.5-32B-Instruct Q4KM: pbhappliedsystems.com/models/qwen2.5-32b-instruct#Q4_K_M
- Qwen2.5-7B-Instruct Q4KM: pbhappliedsystems.com/models/qwen2.5-7b-instruct#Q4_K_M
Models without a published v7.22 run show no scores here; their notes are labelled as v7.21 observations.
Community
Want a pair of models side by side that isn't here yet, or have results of your own to compare?
- This Space — suggest a pair in the Community tab
- Discord — PBH Applied Systems:
#model-requests,#your-evals,#deployment-help - GitHub Discussions — quant_eval Public Corpus tools: model requests and polls on what to measure next
<!-- pbh-community -->
