cloudsurf-software/CloudSurf-4B-FC-GGUF
CloudSurf-4B-FC — official GGUF quants
Official GGUF quantizations of CloudSurf-4B-FC, a Gemma-4 E4B (effective-4B active; ~8.0B total params) function-calling / tool-use model by CloudSurf Software, quantized by us from the BF16 merged champion checkpoint. All benchmark scores, training details, and honesty disclosures live in the base model card — the scores there were measured on the BF16 weights, not on these quants. A quantized-model eval is pending and will be added here when it lands; until then, assume some quality loss below Q8_0, growing as the bit-width shrinks.
Every model file in this repo passed a 6-case tool-call smoke suite (simple call, multi-argument, tool selection, role:tool round-trip, irrelevance refusal, parallel calls) against llama-server --jinja before upload. A file that fails any case does not ship.
Files
Vision (mmproj)
Pair any quant with a projector file for image input: mmproj-CloudSurf-4B-FC.F16.gguf (1.0 GB) or mmproj-CloudSurf-4B-FC.Q8_0.gguf (0.6 GB):
llama-server -m CloudSurf-4B-FC.Q4_K_M.gguf --mmproj mmproj-CloudSurf-4B-FC.F16.gguf --jinja -ngl 99Function calling is this model's primary capability; vision rides along from the Gemma-4 base.
For imatrix (weighted) quants below Q3, see the community repo mradermacher/CloudSurf-4B-FC-i1-GGUF — an independent quantization of the same weights, credit to mradermacher for the fast pickup.
SHA-256 checksums for every file are in `quant-manifest.json`.
Running
llama.cpp (server, OpenAI-compatible tool calls)
--jinja is required — it enables the embedded chat template's tool-call rendering. Without it the server falls back to a generic template and tool calls will not parse.
llama-server -m CloudSurf-4B-FC.Q4_K_M.gguf --jinja -ngl 99 -c 16384 \
--host 127.0.0.1 --port 8080 --alias cloudsurf-4b-fc
# then point any OpenAI client at http://127.0.0.1:8080/v1
# (chat/completions with tools in, tool_calls out)llama.cpp (CLI)
llama-cli -m CloudSurf-4B-FC.Q4_K_M.gguf --jinja -ngl 99 -c 16384 \
-p "You are a helpful assistant." -cnvLM Studio
Search for CloudSurf-4B-FC-GGUF in the in-app model browser and download a quant. Enable the model's own prompt template (default) and use the OpenAI-compatible local server for tool calls.
Ollama
cat > Modelfile <<'EOF'
FROM ./CloudSurf-4B-FC.Q4_K_M.gguf
EOF
ollama create cloudsurf-4b-fc -f Modelfile
ollama run cloudsurf-4b-fcThe GGUF's embedded chat template carries the tool-call format; recent Ollama versions read it directly.
Notes and known quirks
- Scores: the base card's BFCL V4 numbers (3-seed mean 55.73 as-registered / 54.91 matched-variant — both frames disclosed there) are BF16 measurements. No number in that card applies to these quants until the pending quantized-model eval is published here.
- Params: "4B" is effective active parameters (the base's E4B convention); total is ~8.0B, which is what sizes these files.
- `</s>` metadata quirk: GGUF conversions of this model can carry a stray
</s>end-token metadata entry; llama.cpp detects and auto-corrects it at load (benign — observed and verified 2026-08-18). - Context: trained/evaluated at 131K-class context;
-c 16384above is a practical default, raise as your memory allows.
License
Apache-2.0, same as the base model. Gemma-4 base weights are subject to Google's Gemma terms as described in the base model card.
