Team Ai
Datasetpublic

mkvn/quantization-cache-amplification

Quantization as Cache Amplification Trillion-Parameter Mixture-of-Experts Inference on a Commodity Laptop Kavin Kumar, Neural Metrics πŸ“„ Read the paper β€” 11 pages What this is Weight quantization is usually justified as footprint reduction. This work argues that for offloaded mixture-of-experts inference that framing misses the leverage. The binding resource is not storage capacity but the fraction of expert slots resident in DRAM β€” and storage traffic depends on… See the full description on the dataset page: https://huggingface.co/datasets/mkvn/quantization-cache-amplification.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes38downloads
settings

This repository belongs to mkvn on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namequantization-cache-amplification
visibilitypublic
licenceother
gatedno
ownermkvn
Account settings
mkvn/quantization-cache-amplification Β· Team Ai