rubybear/FastContext-1.0-4B-SFT-mlx-8bit
014
FastContext-1.0-4B-SFT-mlx-8bit
8-bit MLX quantization of microsoft/FastContext-1.0-4B-SFT for Apple Silicon.
Quantization details
- Method: Affine 8-bit
- Group size: 64
- Effective bits per weight: 8.5
- Model size: 4.0 GB (vs 7.5 GB bf16)
Benchmark results
Tested on 10 SWE-bench Multilingual instances against other quantization variants:
Highest quality quantization — best File F1 and Line F1 at the cost of larger size and slower inference.
Usage
from mlx_lm import load, generate
model, tokenizer = load("rubybear/FastContext-1.0-4B-SFT-mlx-8bit")Or with fastcontext-mcp for Claude Code integration.
