Team Ai
Datasetpublic

pbhappliedsystems/quant_eval_paired_degradation_statistics

quant_eval — Paired degradation statistics One row per run per task family: the paired pass-rate difference with a 95% confidence interval, the two-sided exact McNemar test, the full discordance breakdown, and a semantic-cluster-adjusted delta and interval. Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing. Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_paired_degradation_statistics.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes58downloads
Dataset Card

quant_eval — Paired degradation statistics

One row per run per task family: the paired pass-rate difference with a 95% confidence interval, the two-sided exact McNemar test, the full discordance breakdown, and a semantic-cluster-adjusted delta and interval.

Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.

Cite this dataset: 10.5281/zenodo.22010557 — concept DOI, always resolves to the latest version. This exact deposit: 10.5281/zenodo.22010558 — version DOI, frozen. Cite this one where reported numbers must stay verifiable against the object referenced.

What this file contains

FileRowsColumns
quant_eval_paired_degradation_statistics.csv4828

Supporting files: source_bundle_checksums.json.

Corpus scope

RunModelBaselineQuantizedSubstrateLicence
Mistral_Nemo_Instruct_2407_20260814_030505mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq4k_mlocalApache-2.0
Mistral_Nemo_Instruct_2407_20260815_113254mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq5k_mlocalApache-2.0
Mistral_Nemo_Instruct_2407_20260816_084553mistralai/Mistral-Nemo-Instruct-2407gguf_f16ggufq80localApache-2.0
Qwen2.5_14B_Instruct_1M_20260815_220633Qwen/Qwen2.5-14B-Instruct-1Mmodal_f16modalq4k_mModalApache-2.0
Qwen2.5_32B_Instruct_20260815_081051Qwen/Qwen2.5-32B-Instructmodal_f16modalq4k_mModalApache-2.0
Qwen2.5_7B_Instruct_20260814_234822Qwen/Qwen2.5-7B-Instructgguf_f16ggufq4k_mlocalApache-2.0

Every run evaluates a full-weight baseline and a quantized variant of the same model against the identical locked fixture set, case for case. Statistical comparison is paired: the two-sided exact McNemar test on per-case outcomes, with Wilson intervals on the rates.

Figures

[image]

Paired pass-rate difference (quantized minus full weight) with 95% confidence intervals, Mistral-Nemo-Instruct-2407, n = 200 paired cases per family. Solid markers indicate two-sided exact McNemar p < 0.05. Families significant: Q4KM 6 of 8; Q5KM 0 of 8; Q8_0 0 of 8. The shape is a discontinuity, not a gradient.

[image]

Paired pass-rate difference at Q4_K_M versus each model's own F16 baseline, n = 200 paired cases per family. Solid markers indicate McNemar p < 0.05. CONFOUNDED: 7B ran on local llama.cpp with a fixed seed; 14B-1M and 32B ran on Modal, which records seed status unsupported. Scale and substrate are not separated by this corpus, and this figure must not be read as a clean parameter-scale effect.

Figures are generated directly from the harness rollups by the published build tooling; no plotted value is recomputed, smoothed, or fitted.

Columns

quant_eval_paired_degradation_statistics.csv

  run_id                                        model_id                                      baseline_quant_type
  quantized_quant_type                          family                                        runner_a
  runner_b                                      paired_n                                      both_pass
  both_fail                                     runner_a_only_pass                            runner_b_only_pass
  disagreements                                 disagreement_rate                             pass_rate_delta_runner_b_minus_runner_a
  delta_ci_low                                  delta_ci_high                                 delta_ci_level
  delta_ci_method                               mcnemar_p_value                               mcnemar_discordant_pairs
  mcnemar_alternative                           discriminating                                cluster_adjusted_method
  cluster_adjusted_effective_paired_n           cluster_adjusted_delta                        cluster_adjusted_delta_ci_low
  cluster_adjusted_delta_ci_high

Verification

This corpus is derived from sanitized publication bundles produced by the quant_eval harness. It is designed to be checked rather than trusted:

  • —source_bundle_checksums.json, included here, republishes, verbatim, the SHA-256 digest and byte length of every file in every source bundle. No source file was modified.
  • —Before this file was written, the builder verified all 72 source-file digests and independently recomputed all 96 family x runner pass rates from the raw per-case rows, matching the harness rollups exactly.
  • —Every row in this file is recomputable from the per-case results dataset (D1) using the gate definitions in data_dictionary.json, published with the per-case results dataset (D1) and the throughput telemetry dataset (D2). The pass counts, rates, and full discordance breakdown follow directly from the per-case outcomes; no intermediate artifact is required.

Limits you should know before using this

  • —Decoding conditions are not uniform across models. Temperature follows each publisher's own model card, so cross-model comparison of absolute pass rates is confounded. Within-run pairing is unaffected, which is what the paired test requires. The conditions are published per row and per run so they can be filtered on.
  • —Runs on the Modal substrate record `seed` status `unsupported` — the deployed method signature accepts no seed — so those runs are not exactly reproducible. Local runs applied a fixed seed.
  • —Runtime figures are observed harness wall time on the recorded hardware and backends. They are not a controlled throughput benchmark and not a general claim about quantization performance at any precision on any hardware. Direction is published as an explicit label because not every measured pair is a speedup.
  • —The fuzz family is an adaptive trajectory evaluated from identical starting fixtures. Its paired test compares complete case outcomes, not identical post-divergence prompts.
  • —Calibration runs are not published. Runs that informed a published run are disclosed by identifier in calibration_lineage.csv, published in the run provenance dataset (D4), so the record is complete without releasing provisional numbers.

Citation

bibtex
@dataset{pbh_quant_eval_d5,
  author    = {Hill, Patrick},
  title     = {quant_eval Paired degradation statistics},
  publisher = {PBH Applied Systems, LLC},
  year      = {2026},
  doi       = {10.5281/zenodo.22010557},
  note      = {Version DOI: 10.5281/zenodo.22010558},
  license   = {CC-BY-4.0}
}

Licence

Creative Commons Attribution 4.0 International (CC BY 4.0). See LICENSE. Commercial use is permitted; attribution is required.

This corpus describes third-party models and redistributes no model weights. Each evaluated model remains under its own licence, recorded per run in the run provenance dataset.


Produced by builddatasets.py 2.4.0 from quanteval publication bundles. Built 2026-08-19.