mkvn/quantization-cache-amplification
Quantization as Cache Amplification Trillion-Parameter Mixture-of-Experts Inference on a Commodity Laptop Kavin Kumar, Neural Metrics π Read the paper β 11 pages What this is Weight quantization is usually justified as footprint reduction. This work argues that for offloaded mixture-of-experts inference that framing misses the leverage. The binding resource is not storage capacity but the fraction of expert slots resident in DRAM β and storage traffic depends onβ¦ See the full description on the dataset page: https://huggingface.co/datasets/mkvn/quantization-cache-amplification.
Quantization as Cache Amplification
Trillion-Parameter Mixture-of-Experts Inference on a Commodity Laptop
Kavin Kumar, Neural Metrics
π [Read the paper](paper/main.pdf) β 11 pages
What this is
Weight quantization is usually justified as footprint reduction. This work argues that for offloaded mixture-of-experts inference that framing misses the leverage. The binding resource is not storage capacity but the fraction of expert slots resident in DRAM β and storage traffic depends on that fraction through a cache hit rate that is both concave and, for recency-based policies, discontinuous.
The central measured result: a least-recently-used expert cache hits exactly zero whenever its capacity falls below the kΒ·L expert slots a single token touches. A token routes to k experts in each of L layers and revisits none of them until the next token β a cyclic reference string, the classical worst case for LRU. On real OLMoE-1B-7B traces (k=8, L=16, so 128 slots) we measure 0.0% hit rate at 2%, 5% and 10% capacity, jumping to 25.3% the moment capacity reaches
- Quantization is what carries a system across that threshold.
Everything here was measured on one laptop: NVIDIA RTX A500 (4 GB VRAM), 32 GB DRAM, consumer NVMe, Windows 11.
Headline numbers
Ablations (all at 1.51 bits, WikiText-2 PPL)
Both codec components are load-bearing, and error feedback matters more than rotation.
Contents
Using the routing traces
import numpy as np
T = np.load("results/routing_trace.npy") # [layers=16, tokens=49152, topk=8]
# distinct expert slots touched by one token:
print(T.shape[0] * T.shape[2]) # 128 -> the LRU thresholdEvery number in the paper is generated programmatically from results/ via code/gen_numbers.py and code/gen_tables.py; nothing is transcribed by hand.
Scope β please read
No trillion-parameter model was executed. No 1T checkpoint was downloaded, quantized, or run. The 1T figures (196 GB at 1.5 bits, 1.81β3.08 tokens/s) are an analytical projection composing measured host parameters with a cache model validated against real 7B-scale routing traces. They are not benchmark results and should not be cited as such. The paper's Limitations section states this, and identifies the weakest assumption: that the Zipf exponent of expert popularity (measured s = 0.65 at 64 experts/layer) is scale-invariant up to 320 experts/layer. A full sensitivity curve across the entire hit-rate range is included precisely because that assumption cannot be foreclosed.
The paper also reports a negative result that constrains any system in this class: sustaining the storage stream while materializing fp16 weights would require ~116 GB/s of device bandwidth against 88.2 GB/s measured, so dequantization must be fused into the GEMM rather than staged through VRAM. No fused kernel was implemented here.
Quality cost is stated plainly rather than buried: sub-2-bit operation on a 1.3B-active-parameter MoE is expensive (8.11 β 22.02 PPL at 1.51 bits), which is why the systems analysis is parameterized by rate rather than asserting a single favourable operating point.
Model used
`allenai/OLMoE-1B-7B-0924` β 6.9B total / 1.3B active, 16 layers, 64 experts/layer, top-8. Evaluation on WikiText-2.
