pbhappliedsystems/quant_eval_paired_degradation_statistics
quant_eval — Paired degradation statistics One row per run per task family: the paired pass-rate difference with a 95% confidence interval, the two-sided exact McNemar test, the full discordance breakdown, and a semantic-cluster-adjusted delta and interval. Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing. Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_paired_degradation_statistics.
quant_eval — Paired degradation statistics
One row per run per task family: the paired pass-rate difference with a 95% confidence interval, the two-sided exact McNemar test, the full discordance breakdown, and a semantic-cluster-adjusted delta and interval.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset: 10.5281/zenodo.22010557 — concept DOI, always resolves to the latest version. This exact deposit: 10.5281/zenodo.22010558 — version DOI, frozen. Cite this one where reported numbers must stay verifiable against the object referenced.
What this file contains
Supporting files: source_bundle_checksums.json.
Corpus scope
Every run evaluates a full-weight baseline and a quantized variant of the same model against the identical locked fixture set, case for case. Statistical comparison is paired: the two-sided exact McNemar test on per-case outcomes, with Wilson intervals on the rates.
Figures
Paired pass-rate difference (quantized minus full weight) with 95% confidence intervals, Mistral-Nemo-Instruct-2407, n = 200 paired cases per family. Solid markers indicate two-sided exact McNemar p < 0.05. Families significant: Q4KM 6 of 8; Q5KM 0 of 8; Q8_0 0 of 8. The shape is a discontinuity, not a gradient.
Paired pass-rate difference at Q4_K_M versus each model's own F16 baseline, n = 200 paired cases per family. Solid markers indicate McNemar p < 0.05. CONFOUNDED: 7B ran on local llama.cpp with a fixed seed; 14B-1M and 32B ran on Modal, which records seed status unsupported. Scale and substrate are not separated by this corpus, and this figure must not be read as a clean parameter-scale effect.
Figures are generated directly from the harness rollups by the published build tooling; no plotted value is recomputed, smoothed, or fitted.
Columns
quant_eval_paired_degradation_statistics.csv
run_id model_id baseline_quant_type
quantized_quant_type family runner_a
runner_b paired_n both_pass
both_fail runner_a_only_pass runner_b_only_pass
disagreements disagreement_rate pass_rate_delta_runner_b_minus_runner_a
delta_ci_low delta_ci_high delta_ci_level
delta_ci_method mcnemar_p_value mcnemar_discordant_pairs
mcnemar_alternative discriminating cluster_adjusted_method
cluster_adjusted_effective_paired_n cluster_adjusted_delta cluster_adjusted_delta_ci_low
cluster_adjusted_delta_ci_highVerification
This corpus is derived from sanitized publication bundles produced by the quant_eval harness. It is designed to be checked rather than trusted:
source_bundle_checksums.json, included here, republishes, verbatim, the SHA-256 digest and byte length of every file in every source bundle. No source file was modified.- Before this file was written, the builder verified all 72 source-file digests and independently recomputed all 96 family x runner pass rates from the raw per-case rows, matching the harness rollups exactly.
- Every row in this file is recomputable from the per-case results dataset (D1) using the gate definitions in
data_dictionary.json, published with the per-case results dataset (D1) and the throughput telemetry dataset (D2). The pass counts, rates, and full discordance breakdown follow directly from the per-case outcomes; no intermediate artifact is required.
Limits you should know before using this
- Decoding conditions are not uniform across models. Temperature follows each publisher's own model card, so cross-model comparison of absolute pass rates is confounded. Within-run pairing is unaffected, which is what the paired test requires. The conditions are published per row and per run so they can be filtered on.
- Runs on the Modal substrate record `seed` status `unsupported` — the deployed method signature accepts no seed — so those runs are not exactly reproducible. Local runs applied a fixed seed.
- Runtime figures are observed harness wall time on the recorded hardware and backends. They are not a controlled throughput benchmark and not a general claim about quantization performance at any precision on any hardware. Direction is published as an explicit label because not every measured pair is a speedup.
- The fuzz family is an adaptive trajectory evaluated from identical starting fixtures. Its paired test compares complete case outcomes, not identical post-divergence prompts.
- Calibration runs are not published. Runs that informed a published run are disclosed by identifier in
calibration_lineage.csv, published in the run provenance dataset (D4), so the record is complete without releasing provisional numbers.
Citation
@dataset{pbh_quant_eval_d5,
author = {Hill, Patrick},
title = {quant_eval Paired degradation statistics},
publisher = {PBH Applied Systems, LLC},
year = {2026},
doi = {10.5281/zenodo.22010557},
note = {Version DOI: 10.5281/zenodo.22010558},
license = {CC-BY-4.0}
}Licence
Creative Commons Attribution 4.0 International (CC BY 4.0). See LICENSE. Commercial use is permitted; attribution is required.
This corpus describes third-party models and redistributes no model weights. Each evaluated model remains under its own licence, recorded per run in the run provenance dataset.
Produced by builddatasets.py 2.4.0 from quanteval publication bundles. Built 2026-08-19.
