Team Ai
Modelpublic

kizuna-intelligence/Qwen3.5-2B-OneCompression-4bit-MLX

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes109downloads
Model Card

Qwen3.5-2B OneCompression 4-bit MLX

Text-only MLX checkpoint produced from Qwen/Qwen3.5-2B for AyaneSDK's on-device conversation example.

  • —OneCompression GPTQ with quantization-error propagation (QEP)
  • —4-bit weights, group size 128
  • —256 Japanese dialogue calibration samples of 512 tokens
  • —186 quantized linear layers
  • —Token embedding quantized separately to asymmetric MLX 4-bit
  • —Vision weights are not included

The packed GPTQ-v1 linear weights were converted losslessly to MLX's row-major affine representation. The model configuration retains an onecompression_source_quantization audit record.

Source

Base model: Qwen/Qwen3.5-2B

Quantizer: FujitsuResearch/OneCompression