Team Ai
Apppublic

yasu-oh/quantized_model_memory_estimator

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes
App README

Quantized model memory estimator

Hugging Face のモデルIDから safetensors.total を取得し、llama.cpp の bpw(bits per weight) で量子化後の重み容量を概算します。

llama.cpp 対応確認は以下の TEXT_MODEL_MAP を参照します。

https://raw.githubusercontent.com/ggml-org/llama.cpp/refs/heads/master/conversion/_init_.py

text
GiB = パラメータ数 × bpw ÷ 8 ÷ 1024^3
python
QUANTIZE = {
    "IQ4_XS": 4.3,
    "Q4_K_M": 4.9,
    "Q5_K_M": 5.7,
    "Q6_K": 6.6,
}

KV cache、CUDA context、allocator fragmentation、temporary workspace、実行時バッファは含みません。