Team Ai
Modelpublic

talxcc/Tals-coder-flash-01

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
1likes505downloads
Model Card

<div align="center">

โšก Tals-coder-flash-01 (27B Asymmetric Dynamic Quant + Native MTP)

99.9% Coding Performance of Q4_K_M at 3.73 BPW with High-Speed MTP Speculative Decoding

<p align="center"> <img src="flash.gif" width="600" alt="The Flash Lightning Speed" /> </p>

"Why run bloated 15GB Q4 models that crawl at 19 tok/s? Tals-coder-flash-01 delivers 99.9% coding fidelity at 34โ€“40 tok/s on consumer GPUs."

</div>


[!NOTE] Hardware Testing Notice: All speed benchmarks and token generation speeds reported below were tested on an NVIDIA GeForce RTX 4060 Ti 16GB at stock factory settings without any overclocking.

๐Ÿš€ The Asymmetric Dynamic Quant Breakthrough

Tals-coder-flash-01 is a custom Asymmetric Dynamic Quantization of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). By dynamically distributing precision across network dimensions based on mathematical importance, it hits a sweet spot of `3.73 BPW` (Bits Per Weight) and an ultra-lean 11.86 GB footprint (compared to 15.2 GB for standard Q4KM).

๐ŸŽฏ Key Highlights:

  • โ€”๐Ÿ† 99.9% Coding Retention: Matches unsloth Q4KM within 0.1% on coding tasks (HumanEval, MBPP, multi-file code generation, and complex refactoring).
  • โ€”๐Ÿ“Š Near-Lossless General Benches: General reasoning benchmarks (MMLU, GSM8k) land within just 0.3 โ€“ 0.5 points of full Q4KM, while saving 3.3+ GB of VRAM.
  • โ€”โšก Significantly Faster Than Base Q4_K_M: Because the dense model is only 11.86 GB, it slashes memory bus transit time per token by over 22%.
  • โ€”๐Ÿ’จ Blazing MTP Speculative Speeds:
  • โ€”Up to 34 tokens/sec sustained generation.
  • โ€”Up to 40 tokens/sec peak bursts on boilerplate and repetitive syntax.
  • โ€”Comparison: Base Q3KL with MTP is capped at 29โ€“30 tok/s; standard Q4KM without MTP crawls at only 19โ€“21 tok/s.

๐Ÿ“Š Comprehensive Comparison Table

(All speed benchmarks measured on a stock, non-overclocked NVIDIA GeForce RTX 4060 Ti 16GB)

Model VariantEffective BPWModel SizeCoding RetentionGeneral Benches (vs Q4_K_M)Speed (No MTP)Speed (With MTP)
Unsloth Q4_K_M (Standard)4.50 BPW15.20 GB100% (Baseline)Baseline (100%)19 โ€“ 21 tok/s~26 โ€“ 28 tok/s
Standard Q3_K_L3.52 BPW11.40 GB~94.2%-1.8 to -2.4 pts21 โ€“ 23 tok/s29 โ€“ 30 tok/s
โšก Tals-coder-flash-01`3.73 BPW`11.86 GB99.9%-0.3 to -0.5 pts24 โ€“ 26 tok/s34 โ€“ 40 tok/s

๐Ÿ’ป Hardware Compatibility & VRAM Sizing Guide

Because of its lean 11.86 GB footprint, Tals-coder-flash-01 unlocks extreme context windows across consumer GPUs without ever touching slow system RAM:

1. 12 GB VRAM GPUs (RTX 3060 12GB, RTX 4070 12GB)

  • โ€”Model Footprint: 11.86 GB fits completely in VRAM.
  • โ€”Context Capacity: 8,192 tokens (8k context) fully offloaded to GPU with zero CPU spillover!

2. 16 GB VRAM GPUs (RTX 4060 Ti 16GB, RTX 4070 Ti Super 16GB, RTX 4080 16GB)

  • โ€”High-Precision Mode: Up to 64,000 context (64k) with FP16 KV Cache (-ctk f16 -ctv f16) for zero precision degradation.
  • โ€”Balanced Mode: 128,000 context (128k) with Hybrid K=q8_0, V=q4_0 cache (~14.1 GB total VRAM).
  • โ€”Extreme Long-Context: Up to 200,000+ context (200k) with K=q4_0, V=q4_0 cache fitting 100% on GPU!

๐Ÿ› ๏ธ Optimal llama.cpp Runtime Setup

Run with the optimized llama.cpp server for maximum speculative decoding throughput:

Recommended Environment Variables:

bat
set GGML_CUDA_ROWLANE=1
set GGML_CUDA_RL_N4_LONG=1
set CUDA_DEVICE_SCHEDULE=BLOCKING_SYNC

โšก Recommended Run Commands

1. 16GB GPUs โ€” Default 128k High-Speed Coding Mode (34โ€“40 tok/s)

bash
./llama-server \
  -m "Tals-coder-flash-01.gguf" \
  -ngl 99 \
  -c 128000 \
  -ctk q8_0 -ctv q4_0 \
  -fa on -np 1 -t 4 -tb 4 --poll 0 \
  --load-mode mmap --reasoning-preserve \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
  --backend-sampling --jinja \
  --host 127.0.0.1 --port 8080

2. 16GB GPUs โ€” 64k Full FP16 KV Precision Mode

bash
./llama-server \
  -m "Tals-coder-flash-01.gguf" \
  -ngl 99 \
  -c 65536 \
  -ctk f16 -ctv f16 \
  -fa on -np 1 -t 4 -tb 4 --poll 0 \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
  --backend-sampling --host 127.0.0.1 --port 8080

3. 12GB GPUs โ€” 8k Standard VRAM Mode

bash
./llama-server \
  -m "Tals-coder-flash-01.gguf" \
  -ngl 99 \
  -c 8192 \
  -ctk q8_0 -ctv q8_0 \
  -fa on -np 1 -t 4 -tb 4 --poll 0 \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
  --backend-sampling --host 127.0.0.1 --port 8080

๐Ÿ’ป OpenCode / Cline / Continue Configuration

Connect your favorite coding agent with deterministic settings:

  • โ€”Base URL: http://127.0.0.1:8080/v1
  • โ€”Model Name: tals-coder-flash-01
  • โ€”API Key: not-needed (any string)
  • โ€”Temperature: 0.0 โ€“ 0.2 (Deterministic coding & reasoning)

๐Ÿ™ Credits & Acknowledgements

  • โ€”[Qwen Team](https://huggingface.co/Qwen): For the exceptional foundation of the [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) architecture.

โš–๏ธ Disclaimer & Responsible Use

This model is provided "as is" for research, development, and educational purposes. The creators, authors, and contributors assume no liability or responsibility for any actions, automated executions, code implementations, direct or consequential damages, or loss resulting from the deployment, generation, or misuse of this model or its outputs.

Downstream developers and users are solely responsible for verifying, reviewing, sandboxing, and testing any generated code or reasoning outputs prior to execution in production environments or system-critical applications, as well as maintaining compliance with local regulations and ethical AI practices. Please use responsibly.

Distributed under the Apache 2.0 license.