Team Ai
Modelpublic

SherlockID365/Qwen3-VL-8B-Instruct-quantized.w4a16

sourceHugging Faceapache-2.0updated 11mo agoView on Hugging Face
1likes71kdownloads
Model Card

code taken from : https://github.com/vllm-project/llm-compressor/blob/main/examples/awq/qwen3-vl-30b-a3b-Instruct-example.py

Qwen3-VL-8B-Instruct-AWQ

AWQ (W4A16) quantized version of Qwen/Qwen3-VL-8B-Instruct.

  • —Quantization: AWQ, 4 bits, groupsize=128, zeropoint=true, version="gemm"
  • —modules_to_not_convert: ["visual"]
  • —Prepared with LLM Compressor oneshot AWQ. recipe = AWQModifier( targets="Linear", scheme="W4A16", ignore=[r"re:model.visual.", r"re:visual."], # drop lmhead from ignore duoscaling=True, )