mkvn/quantization-cache-amplification
Quantization as Cache Amplification Trillion-Parameter Mixture-of-Experts Inference on a Commodity Laptop Kavin Kumar, Neural Metrics π Read the paper β 11 pages What this is Weight quantization is usually justified as footprint reduction. This work argues that for offloaded mixture-of-experts inference that framing misses the leverage. The binding resource is not storage capacity but the fraction of expert slots resident in DRAM β and storage traffic depends onβ¦ See the full description on the dataset page: https://huggingface.co/datasets/mkvn/quantization-cache-amplification.
049
Update dataset caption
Paper, codec, routing traces and measurements
initial commit
