talxcc/Tals-coder-flash-02
<div align="center">
β‘ Tals-coder-flash-02 (27B Ternary + Grafted Native MTP + Multimodal Vision)
The Fastest 27B Coding Model on Consumer Hardware
<p align="center"> <img src="flash.gif" width="600" alt="The Flash Lightning Speed" /> </p>
"Why wait 5 minutes for a reasoning model to finish an existential crisis in its chain-of-thought? Tals-coder-flash-02 goes at lightning speed with 60 tok/s on a budget 16GB GPU."
<table> <tr> <td width="50%"><img src="proof.gif" width="100%" alt="Proof" /></td> <td width="50%"><img src="metrics.gif" width="100%" alt="Metrics" /></td> </tr> </table> </div>
[!NOTE] Hardware Testing Notice: All speed benchmarks and token generation speeds reported below were tested on an NVIDIA GeForce RTX 4060 Ti 16GB at stock factory settings without any overclocking temp 0.2 top p 0.85.
ποΈ Pure Speed: Breaking the Physical Bandwidth Limit
On paper, an NVIDIA RTX 4060 Ti 16GB has a narrow 128-bit memory bus with 288 GB/s bandwidth. Streaming an 8 GB 27B model on that bus physically caps standard autoregressive generation to ~25 tokens/sec.
Tals-coder-flash-02 shatters this barrier:
- β‘ 55β60 tokens/second sustained on RTX 4060 Ti 16GB (verified in production!).
- π Up to 76.7 tokens/second peak burst during boilerplate and repetitive code blocks.
- π¨ 120β150+ tokens/second projected on high-bandwidth hardware (RTX 3090, 4090, A100).
- π― Zero Overthinking: No rambling monologues, no 2,000-token loops questioning itself. It cuts straight to clean, functional code.
𧬠Custom GGUF Quantization & Grafting
Tals-coder-flash-02 is a custom GGUF quantization of [prism-ml/Ternary-Bonsai-2-27B-gguf](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf):
- Surgically Grafted Native MTP: We permanently grafted the 65th Multi-Token Prediction (MTP) draft head and unrotated token embeddings directly into the main weights (
Tals-coder-flash-02.gguf). You do not need to juggle separate target and drafter model filesβit's an all-in-one standalone file with native speculative decoding! - Multimodal Vision Projector: Includes the high-fidelity
Tals-coder-flash-02-mmproj.ggufCLIP Q8_0 vision tower (0.63 GB) for analyzing UI designs, charts, diagrams, and debugging screenshots. - 100% VRAM Execution: Runs 27B parameters, 128,000 context, and multimodal vision completely on GPU (~13.8β14.1 GB / 16.0 GB total) with zero CPU/PCIe spillover.
π Benchmark & Performance
Tested on NVIDIA GeForce RTX 4060 Ti 16GB (PCIe 4.0 x8, 288 GB/s bandwidth, stock factory settings without overclocking):
π¦ Model Files in Repository
π οΈ Optimal llama.cpp Runtime Setup
Because this model uses ternary quantization in a Hadamard-rotated basis, use the optimized llama.cpp build with Hadamard kernel support (Prism ML / Pascal branch, or apply bonsai2-pascal.patch).
Critical Environment Variables (Mandatory)
Before starting the server, set these environment variables to enable Pascal CUDA kernels and prevent CPU busy-polling:
Windows (PowerShell / CMD):
set GGML_CUDA_ROWLANE=1
set GGML_CUDA_RL_N4_LONG=1
set CUDA_DEVICE_SCHEDULE=BLOCKING_SYNCLinux (Bash):
export GGML_CUDA_ROWLANE=1
export GGML_CUDA_RL_N4_LONG=1
export CUDA_DEVICE_SCHEDULE=BLOCKING_SYNCβ‘ Recommended Run Configurations
1. Default Production Mode (128k Context + Grafted MTP + Vision)
Run the single unified model with vision on a 16GB GPU:
./llama-server \
-m "Tals-coder-flash-02.gguf" \
--mmproj "Tals-coder-flash-02-mmproj.gguf" \
--mmproj-offload \
--image-min-tokens 1024 \
-ngl 99 \
-c 128000 \
-ctk q8_0 -ctv q4_0 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--load-mode mmap --reasoning-preserve \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --jinja \
--host 127.0.0.1 --port 80802. High-Precision Coding Mode (64k Context + FP16 KV Cache)
Maximum precision for complex mathematical derivations:
./llama-server \
-m "Tals-coder-flash-02.gguf" \
--mmproj "Tals-coder-flash-02-mmproj.gguf" \
-ngl 99 \
-c 65536 \
-ctk f16 -ctv f16 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --host 127.0.0.1 --port 80803. Extreme Context Mode (262k Native Context on 16GB VRAM)
Process entire codebases and large book-length documents with zero CPU offload:
./llama-server \
-m "Tals-coder-flash-02.gguf" \
-ngl 99 \
-c 262144 \
-ctk q4_0 -ctv q4_0 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --host 127.0.0.1 --port 8080π» OpenCode / Cline / Continue Configuration
Set your tool's API endpoint to:
- Base URL:
http://127.0.0.1:8080/v1 - Model Name:
tals-coder-flash-02 - API Key:
not-needed(any string) - Temperature:
0.0β0.2(Deterministic coding & reasoning)
π Credits & Acknowledgements
- [PrismML](https://huggingface.co/prism-ml): The original creators of the [Ternary-Bonsai-2-27B-gguf](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) architecture and pioneering 1.58-bit ternary Hadamard quantization kernels.
- [Ukisai](https://huggingface.co/ukisai/Swift-Bonsai-2-GGUF): For Swift-Bonsai-2, adapting and quantizing the ternary base weights.
- [killy369 / Kilian](https://huggingface.co/killy369/Ternary-Bonsai-2-27B-MTP-drafter-GGUF): For training and providing the native MTP speculative draft weights and 64k pruned vocabulary.
- Qwen Team: For the underlying Qwen 3.5 architecture.
βοΈ Disclaimer & Responsible Use
This model is provided "as is" for research, development, and educational purposes. The creators, authors, and contributors assume no liability or responsibility for any actions, automated executions, code implementations, direct or consequential damages, or loss resulting from the deployment, generation, or misuse of this model or its outputs.
Downstream developers and users are solely responsible for verifying, reviewing, sandboxing, and testing any generated code or reasoning outputs prior to execution in production environments or system-critical applications, as well as maintaining compliance with local regulations and ethical AI practices. Please use responsibly.
Distributed under the Apache 2.0 license.
