leocyusa/ra-quant-vulnerability-maps
Reasoning-Aware Quantization: Vulnerability Maps for DeepSeek-R1 Distilled Models Pre-computed, per-module INT4 vulnerability maps for DeepSeek-R1 distilled reasoning models, plus accuracy–energy results for selective mixed-precision compression. Use the maps to decide which layers to keep in higher precision without repeating the profiling (336 single-module evaluations per benchmark for a 14B model). Produced with compute provided by OpenToken, at Carnegie Mellon University… See the full description on the dataset page: https://huggingface.co/datasets/leocyusa/ra-quant-vulnerability-maps.
Reasoning-Aware Quantization: Vulnerability Maps for DeepSeek-R1 Distilled Models
Pre-computed, per-module INT4 vulnerability maps for DeepSeek-R1 distilled reasoning models, plus accuracy–energy results for selective mixed-precision compression. Use the maps to decide which layers to keep in higher precision without repeating the profiling (336 single-module evaluations per benchmark for a 14B model).
Produced with compute provided by OpenToken, at Carnegie Mellon University Africa (Leonard Twagirayezu; supervisor Prof. Prasenjit Mitra). Code: https://github.com/Leonard2019-samantha/ra_quant
Contents
Benchmarks: GSM8K, MATH-500, ProofWriter (depth-5), FOLIO, MuSiQue (token F1).
Vulnerability map columns
Method
- Quantization: BitsAndBytes NF4, weight-only, round-to-nearest, float16 compute.
- Sweep: with the FP16 model resident, each (layer, projection) pair is quantized alone, the calibration split (first 50% of the dataset) is evaluated, and the pair is restored. 14B: 15 calibration questions per benchmark; 7B: 30.
- Selective results: evaluated on the held-out second 50%; 14B: 30 questions × 3 runs.
- Energy: NVML power sampled every 100 ms; energy = mean power × wall-clock time.
- Hardware: 1× Tesla V100-SXM3 32 GB (350 W limit), OpenToken.
Usage
import pandas as pd
m = pd.read_csv("vulnerability_maps/qwen14b_gsm8k.csv")
k = int(0.10 * len(m)) # protect the top 10%
keep_fp16 = m.nsmallest(k, "rank")["full_name"].tolist()
# Pass keep_fp16 as llm_int8_skip_modules in BitsAndBytesConfig,
# or restore those modules from an FP16 copy after loading in 4-bit.Caveats
- Many modules tie on
accuracy_dropbecause calibration sets are small; treat ranks within a tie as equal. - MuSiQue calibration accuracy is near zero (14B: 1 of 15), so its map is not informative.
- In the 14B selective runs, non-quantized layers and restored modules ran in bfloat16 (FP16 baseline in float16); V100s lack bfloat16 tensor cores, which may raise the measured energy of INT4 and top-k conditions.
- Energy values are specific to the V100-SXM3 and are not comparable across GPUs.
Citation
Twagirayezu, L. and Mitra, P. (2026). Reasoning-Aware Compression: Identifying and
Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment.
Under review, ACL Rolling Review.Acknowledgement and license
Compute for these maps was provided by OpenToken. Released under CC BY 4.0: free to use and adapt with attribution.
