talxcc/Tals-coder-flash-01
<div align="center">
โก Tals-coder-flash-01 (27B Asymmetric Dynamic Quant + Native MTP)
99.9% Coding Performance of Q4_K_M at 3.73 BPW with High-Speed MTP Speculative Decoding
<p align="center"> <img src="flash.gif" width="600" alt="The Flash Lightning Speed" /> </p>
"Why run bloated 15GB Q4 models that crawl at 19 tok/s? Tals-coder-flash-01 delivers 99.9% coding fidelity at 34โ40 tok/s on consumer GPUs."
</div>
[!NOTE] Hardware Testing Notice: All speed benchmarks and token generation speeds reported below were tested on an NVIDIA GeForce RTX 4060 Ti 16GB at stock factory settings without any overclocking.
๐ The Asymmetric Dynamic Quant Breakthrough
Tals-coder-flash-01 is a custom Asymmetric Dynamic Quantization of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). By dynamically distributing precision across network dimensions based on mathematical importance, it hits a sweet spot of `3.73 BPW` (Bits Per Weight) and an ultra-lean 11.86 GB footprint (compared to 15.2 GB for standard Q4KM).
๐ฏ Key Highlights:
- ๐ 99.9% Coding Retention: Matches unsloth Q4KM within 0.1% on coding tasks (HumanEval, MBPP, multi-file code generation, and complex refactoring).
- ๐ Near-Lossless General Benches: General reasoning benchmarks (MMLU, GSM8k) land within just 0.3 โ 0.5 points of full Q4KM, while saving 3.3+ GB of VRAM.
- โก Significantly Faster Than Base Q4_K_M: Because the dense model is only 11.86 GB, it slashes memory bus transit time per token by over 22%.
- ๐จ Blazing MTP Speculative Speeds:
- Up to 34 tokens/sec sustained generation.
- Up to 40 tokens/sec peak bursts on boilerplate and repetitive syntax.
- Comparison: Base Q3KL with MTP is capped at 29โ30 tok/s; standard Q4KM without MTP crawls at only 19โ21 tok/s.
๐ Comprehensive Comparison Table
(All speed benchmarks measured on a stock, non-overclocked NVIDIA GeForce RTX 4060 Ti 16GB)
๐ป Hardware Compatibility & VRAM Sizing Guide
Because of its lean 11.86 GB footprint, Tals-coder-flash-01 unlocks extreme context windows across consumer GPUs without ever touching slow system RAM:
1. 12 GB VRAM GPUs (RTX 3060 12GB, RTX 4070 12GB)
- Model Footprint: 11.86 GB fits completely in VRAM.
- Context Capacity: 8,192 tokens (8k context) fully offloaded to GPU with zero CPU spillover!
2. 16 GB VRAM GPUs (RTX 4060 Ti 16GB, RTX 4070 Ti Super 16GB, RTX 4080 16GB)
- High-Precision Mode: Up to 64,000 context (64k) with FP16 KV Cache (
-ctk f16 -ctv f16) for zero precision degradation. - Balanced Mode: 128,000 context (128k) with Hybrid
K=q8_0,V=q4_0cache (~14.1 GB total VRAM). - Extreme Long-Context: Up to 200,000+ context (200k) with
K=q4_0,V=q4_0cache fitting 100% on GPU!
๐ ๏ธ Optimal llama.cpp Runtime Setup
Run with the optimized llama.cpp server for maximum speculative decoding throughput:
Recommended Environment Variables:
set GGML_CUDA_ROWLANE=1
set GGML_CUDA_RL_N4_LONG=1
set CUDA_DEVICE_SCHEDULE=BLOCKING_SYNCโก Recommended Run Commands
1. 16GB GPUs โ Default 128k High-Speed Coding Mode (34โ40 tok/s)
./llama-server \
-m "Tals-coder-flash-01.gguf" \
-ngl 99 \
-c 128000 \
-ctk q8_0 -ctv q4_0 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--load-mode mmap --reasoning-preserve \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --jinja \
--host 127.0.0.1 --port 80802. 16GB GPUs โ 64k Full FP16 KV Precision Mode
./llama-server \
-m "Tals-coder-flash-01.gguf" \
-ngl 99 \
-c 65536 \
-ctk f16 -ctv f16 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --host 127.0.0.1 --port 80803. 12GB GPUs โ 8k Standard VRAM Mode
./llama-server \
-m "Tals-coder-flash-01.gguf" \
-ngl 99 \
-c 8192 \
-ctk q8_0 -ctv q8_0 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --host 127.0.0.1 --port 8080๐ป OpenCode / Cline / Continue Configuration
Connect your favorite coding agent with deterministic settings:
- Base URL:
http://127.0.0.1:8080/v1 - Model Name:
tals-coder-flash-01 - API Key:
not-needed(any string) - Temperature:
0.0โ0.2(Deterministic coding & reasoning)
๐ Credits & Acknowledgements
- [Qwen Team](https://huggingface.co/Qwen): For the exceptional foundation of the [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) architecture.
โ๏ธ Disclaimer & Responsible Use
This model is provided "as is" for research, development, and educational purposes. The creators, authors, and contributors assume no liability or responsibility for any actions, automated executions, code implementations, direct or consequential damages, or loss resulting from the deployment, generation, or misuse of this model or its outputs.
Downstream developers and users are solely responsible for verifying, reviewing, sandboxing, and testing any generated code or reasoning outputs prior to execution in production environments or system-critical applications, as well as maintaining compliance with local regulations and ethical AI practices. Please use responsibly.
Distributed under the Apache 2.0 license.
