ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF
Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-GGUF
GGUF quants of ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP, an abliterated Swift 1.5 Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). If you have a GPU with Blackwell architecture: NVFP4 (VLLM/SGLANg).
- Uncensored: 23/100 refusals against Swift 1.5's 98/100, at a KL divergence of 0.0884 from Swift 1.5.
- Two layouts:
UD-*files follow Unsloth's Dynamic 3.0 per-tensor types;RCO-*files, in theRCO/folder, follow ISTA-DASLab's GSQ-RCO allocation (see RCO quants). The RCO files were added on request and are more experimental. - MTP head included in every main GGUF, for self-speculative decoding in llama.cpp.
- Vision: image and video input through the included projector.
Quants
<table> <thead> <tr><th align="left" width="213">Quant</th><th align="right" width="160">Size</th><th align="right" width="160">Mean KLD</th><th align="right" width="160">Same top-1</th></tr> </thead> <tbody> <tr><td>BF16</td><td align="right">50.9 GiB</td><td align="right">—</td><td align="right">—</td></tr> <tr><td>UD-Q8KXL</td><td align="right">29.3 GiB</td><td align="right">0.0068</td><td align="right">98.8 %</td></tr> <tr><td>UD-Q6KXL</td><td align="right">23.6 GiB</td><td align="right">0.0138</td><td align="right">98.2 %</td></tr> <tr><td>UD-Q5KM</td><td align="right">18.4 GiB</td><td align="right">0.0238</td><td align="right">96.8 %</td></tr> <tr><td>UD-Q5KS</td><td align="right">17.4 GiB</td><td align="right">0.0279</td><td align="right">96.4 %</td></tr> <tr><td>UD-Q4KXL</td><td align="right">16.4 GiB</td><td align="right">0.0369</td><td align="right">95.6 %</td></tr> <tr><td>UD-Q4KS</td><td align="right">14.3 GiB</td><td align="right">0.0575</td><td align="right">94.0 %</td></tr> <tr><td>UD-IQ4XS</td><td align="right">13.3 GiB</td><td align="right">0.0691</td><td align="right">92.9 %</td></tr> <tr><td>UD-Q3KXL</td><td align="right">12.2 GiB</td><td align="right">0.0789</td><td align="right">91.7 %</td></tr> <tr><td>RCO-IQ3S</td><td align="right">11.3 GiB</td><td align="right">0.1084</td><td align="right">89.5 %</td></tr> <tr><td>UD-IQ3S</td><td align="right">11.2 GiB</td><td align="right">0.1074</td><td align="right">89.9 %</td></tr> <tr><td>UD-IQ3XXS</td><td align="right">10.2 GiB</td><td align="right">0.1353</td><td align="right">87.8 %</td></tr> <tr><td>RCO-IQ3XXS</td><td align="right">9.7 GiB</td><td align="right">0.1774</td><td align="right">86.4 %</td></tr> <tr><td>UD-Q2KXL</td><td align="right">9.2 GiB</td><td align="right">0.1902</td><td align="right">85.1 %</td></tr> <tr><td>RCO-IQ2S</td><td align="right">8.9 GiB</td><td align="right">0.2349</td><td align="right">83.8 %</td></tr> <tr><td>RCO-IQ2_XS</td><td align="right">8.2 GiB</td><td align="right">0.3018</td><td align="right">81.2 %</td></tr> </tbody> </table>
Files: Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-<quant>.gguf, with the RCO quants under RCO/; the BF16 is Swift-1.5-Qwen3.8-27B-Uncensored-MTP-BF16.gguf. Also here:
Usage (llama.cpp)
# text only
llama-server -m Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-UD-Q6_K_XL.gguf -ngl 99 -c 32768
# with vision (image and video input)
llama-server -m Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-UD-Q6_K_XL.gguf -ngl 99 -c 32768 \
--mmproj mmproj-BF16.gguf
# self-speculative decoding with the built-in MTP head
llama-server -m Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-UD-Q6_K_XL.gguf -ngl 99 -c 32768 \
--spec-type draft-mtp--spec-type draft-mtp needs a llama.cpp build with MTP support for qwen35. The MTP head loads from the main GGUF; there is no separate draft file. With llama-server -hf, name a quant with its prefix (:RCO-IQ3_XXS), since a bare :IQ3_XXS matches both layouts.
Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, min_p 0. The model thinks before answering by default.
Quality
Each quant compared with the BF16 it was made from, token by token: wikitext-2 test set (English prose), 36 × 8192 tokens, f16 KV cache, llama.cpp 94659b076. Mean KLD is the KL divergence of the quant's next-token distribution from the BF16's; same top-1 is the share of tokens where both rank the same token first. Both are relative to this BF16 and this text, so they rank the files here against each other but are not comparable with other repos' figures or with the uncensoring's 0.0884, which is measured against Swift 1.5.
- UD quants: quality follows file size with no exceptions. Above UD-Q5KM a GiB buys little; below it each GiB saved costs about twice as much.
- RCO quants: RCO-IQ3S matches UD-IQ3S. The three smaller ones are further from the BF16 than the UD files around their size. They skip GSQ refinement (see RCO quants), so these numbers do not describe UkisAI's GSQ-RCO files.
- Where the error is: the typical token is barely touched (median KLD 0.0002 at UD-Q8KXL, 0.003 at UD-Q4KXL); the mean is carried by a small share of tokens that change a lot. From UD-Q3KXL down the error also turns one-sided, taking 0.4–1.3 % of probability from the BF16's token on average.
kld_results.csv has every statistic, including the tail percentiles and perplexity. Perplexity does not rank these files: 12 of the 15 quants score below the BF16.
The model
The source applies the refusal direction of orcarouter/Qwen3.8-27B-Uncensored (Arditi et al. 2024, one direction) to Swift 1.5's weights: 131 tensors edited (attention outputs, mlp.down_proj, embed_tokens, MTP layer included), everything else Swift 1.5's. Method, direction and scripts are on the source model's card.
Measured with Heretic on the BF16 weights (100 prompts from mlabonne/harmful_behaviors, keyword refusal detector; first-token KL on mlabonne/harmless_alpaca; thinking skipped). Not re-measured on the quants.
Quantization
- Conversion: the source model's BF16 safetensors (all 1199 tensors, MTP included) were converted with llama.cpp's
convert_hf_to_gguf.py, once for the language model and once with--mmprojfor the vision projector. - UD quants:
llama-quantizewith Unsloth'simatrix_unsloth.gguffrom unsloth/Qwen3.8-27B-GGUF, plus a--tensor-type-filethat sets every tensor's type to the one in Unsloth's GGUF of the same size.
The imatrix was calibrated on the original Qwen3.8-27B; Swift 1.5's fine-tune and the rank-one abliteration keep the architecture and every tensor name, and change the weights only slightly.
RCO quants
The four RCO-* files give each tensor the type chosen by ISTA-DASLab's RCO search for Qwen3.8-27B, as published with ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF and reused for Swift 1.5 in ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF.
These are not UkisAI's GSQ-RCO files made uncensored. The layout and the importance matrix are the same as theirs, but the weights are rounded by llama-quantize as usual, without GSQ's trained refinement step. Sizes match UkisAI's -mtp files to within a few hundred header bytes.
- Allocation: the per-tensor types of UkisAI's
-mtpfiles, read from their publishedtensor-allocation/*-mtp.rco-allocation.txt(commitd74895bb). These are ISTA-DASLab's allocations, plus Q6_K for the MTP head. - Importance matrix: UkisAI's
imatrix-swift15-v1mix.gguffrom the same repository, calibrated on Swift 1.5.
The 96 BF16 tensors are the small ssm_alpha and ssm_beta projections, kept unquantized as in the allocation. To reproduce, build a map from rco_types.tsv as in Reproduce and quantize with --imatrix imatrix-swift15-v1mix.gguf; the fallback types are IQ2XS, IQ2S, IQ3XXS and IQ3S, and every tensor is in the map.
Checks
- Layout: each quant's 866 tensors have exactly the types and shapes of Unsloth's file of the same size; each RCO quant's have exactly the types in UkisAI's allocation dump of the same size.
- MTP: every main GGUF has 65 blocks, with the MTP layer's 15 tensors in
blk.64(nextn_predict_layers = 1). llama.cpp's--spec-type draft-mtploads it and drafts with it. - Tokenizer and chat template: vocabulary, merges and special tokens as in Unsloth's Qwen3.8-27B GGUF;
add_bos_token = falseand the padding token from Swift 1.5's tokenizer config; the chat template byte-identical to Swift 1.5's (Qwen3.8-27B's). - Vision projector: 334 tensors, converted from Swift 1.5's vision tower.
- Load test: every quant loads and runs on a 32 GB GPU: fully offloaded up to UD-Q6KXL, with 56 of 65 layers offloaded for UD-Q8KXL. Each produced sane logits on a short perplexity run before upload; the measurement that ranks them is under Quality.
Reproduce
# 1. convert (Swift-1.5-Qwen3.8-27B-Uncensored-MTP/ is a download of the source repository)
python convert_hf_to_gguf.py Swift-1.5-Qwen3.8-27B-Uncensored-MTP --outtype bf16 \
--model-name Swift-1.5-Qwen3.8-27B-Uncensored-MTP --outfile Swift-1.5-Qwen3.8-27B-Uncensored-MTP-BF16.gguf
python convert_hf_to_gguf.py Swift-1.5-Qwen3.8-27B-Uncensored-MTP --outtype bf16 --mmproj \
--model-name Swift-1.5-Qwen3.8-27B-Uncensored-MTP --outfile mmproj-BF16.gguf
# 2. per-size --tensor-type-file from tensor_types.tsv (here UD-Q5_K_M), one anchored regex per tensor
awk -F'\t' 'NR==1 { for (i=2; i<=NF; i++) if ($i=="UD-Q5_K_M") c=i; next }
{ n=$1; gsub(/\./,"\\\\.",n); print "^" n "$=" $c }' \
tensor_types.tsv > tensor_types_UD-Q5_K_M.txt
# 3. quantize. The last argument is the fallback type for tensors the file doesn't list:
# UD-Q2_K_XL Q2_K, UD-IQ3_XXS IQ3_XXS, UD-IQ3_S IQ3_S, UD-Q3_K_XL Q3_K_L,
# UD-IQ4_XS IQ4_XS, UD-Q4_K_S Q4_K_S, UD-Q4_K_XL Q4_K_M, UD-Q5_K_S Q5_K_S,
# UD-Q5_K_M Q5_K_M, UD-Q6_K_XL Q6_K, UD-Q8_K_XL Q8_0
llama-quantize --imatrix imatrix_unsloth.gguf \
--tensor-type-file tensor_types_UD-Q5_K_M.txt \
Swift-1.5-Qwen3.8-27B-Uncensored-MTP-BF16.gguf \
Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-UD-Q5_K_M.gguf Q5_K_Mtensor_types.tsv holds the type of every non-F32 tensor in each of Unsloth's sizes, one row per tensor and one column per size. The regexes must be anchored (^…$). llama-quantize matches them unanchored, so a bare output\.weight line would also match every attn_output.weight.
Not evaluated
General benchmarks, refusal behaviour in thinking mode and on the quants, whether Swift 1.5's shorter reasoning traces survive, and MTP acceptance against Swift 1.5.
License
Derivative of Swift 1.5 Qwen3.8-27B, under the Swift Open License v1.0 (license): free for individuals and organizations with gross annual revenue up to US$1,000,000; above that, commercial use needs a Swift Enterprise License from UkisAI. Qwen3.8-27B and orcarouter/Qwen3.8-27B-Uncensored are Apache 2.0.
Intended use
The model answers requests the original declines. You are responsible for how you use it and for complying with applicable law and the license.
Credits
- Qwen for Qwen3.8-27B.
- UkisAI for Swift 1.5 Qwen3.8-27B.
- OrcaRouter for Qwen3.8-27B-Uncensored and its refusal direction.
- Unsloth for the Dynamic 3.0 quantization layout and imatrix.
- ISTA-DASLab for the RCO per-tensor allocations (GSQ, RCO), and UkisAI for publishing them for Swift 1.5 together with its importance matrix.
- Arditi et al., Refusal in Language Models Is Mediated by a Single Direction (2024).
- Heretic, used for evaluation.
