dflash2
Qwen3.8-27B-DFlash2-GGUFQwen3.8-27B-DFlash2Qwen3.8-27B-DFlash2Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUFGLM-5.3-Flash-DFlash2Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-DFlash2-NInfer-v3
Datasets
All datasets matching “dflash2”bonsai2-dflash2-bench
Ternary Bonsai 2 27B + DFlash2 on one NVIDIA L4 — benchmark, 2026-09-24
Engine: llama-server from PrismML-Eng/llama.cpp prism @ ee8ad0ef6 + the DFlash2 cherry-pick (PrismML-Eng/llama.cpp#261),
CUDA 12.8 (sm_89). GPU: NVIDIA L4 24 GB (AWS g6.12xlarge, one server per GPU), driver 535.
Greedy (temperature 0), one request at a time, -np 1. Target GGUF prism-ml/Ternary-Bonsai-2-27B-gguf.
Drafter r3 = naklitechie/Qwen3.8-27B-DFlash2-ternary-bonsai2 Q4_K_M; stock =… See the full description on the dataset page: https://huggingface.co/datasets/naklitechie/bonsai2-dflash2-bench.qwen3.8-27b-uncensored-dflash2-m4-pro-benchmark
Qwen3.8-27B (MLX 4-bit) + DFlash2 speculative decoding on M4 Pro — benchmark recipe
This is a benchmark recipe, not redistributed weights. It records the exact
hardware, software, and commands used to measure a 2.06x generation-throughput
speedup with DFlash speculative decoding, and how to rerun it.
Result
HumanEval, 20 samples, max 256 new tokens, temperature 0 (greedy), reasoning
off, block size 5, paired baseline and DFlash under identical settings. Other… See the full description on the dataset page: https://huggingface.co/datasets/hamiejuice/qwen3.8-27b-uncensored-dflash2-m4-pro-benchmark.llama-cpp-dflash2-cuda-bin
