yasu-oh/quantized_model_memory_estimator
0
Quantized model memory estimator
Hugging Face のモデルIDから safetensors.total を取得し、llama.cpp の bpw(bits per weight) で量子化後の重み容量を概算します。
llama.cpp 対応確認は以下の TEXT_MODEL_MAP を参照します。
https://raw.githubusercontent.com/ggml-org/llama.cpp/refs/heads/master/conversion/_init_.py
GiB = パラメータ数 × bpw ÷ 8 ÷ 1024^3QUANTIZE = {
"IQ4_XS": 4.3,
"Q4_K_M": 4.9,
"Q5_K_M": 5.7,
"Q6_K": 6.6,
}KV cache、CUDA context、allocator fragmentation、temporary workspace、実行時バッファは含みません。
