FINAL-Bench/POCKET-35B-GGUF
### π [POCKET-Qwen3.8-Flash-Next](https://huggingface.co/FINAL-Bench/POCKET-Qwen3.8-Flash-Next-GGUF) β a 180B model running on a laptop with 8 GB VRAM + 32 GB RAM Β· 4.17 tok/s measured.  ![VRAM]() ![RAM]() ![Speed]() <!-- POCKET-FLASHNEXT-BADGE -->
### π [POCKET-Zimage-CPU](https://huggingface.co/FINAL-Bench/POCKET-Zimage-CPU) β photoreal images in 46 s on a CPU only. No GPU, no CUDA, no Python.   ![RAM]() <!-- POCKET-Zimage-CPU-BADGE -->
### π Collections βΆ [POCKET Models](https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6) β this family (on-device, no GPU) Darwin Family Β· Aether Foundation Β· VKAE Accelerated
POCKET-35B-GGUF
A 35B model that runs on your PC with no GPU β and on your phone. Just stock llama.cpp. No fork, no CUDA, no cloud.
π Try it live, no install β   β both answering on a CPU-only box (no GPU). POCKET-26B is Gemma4-based.
  ![No GPU]() ![Base]()
Pick your build β        
The POCKET lineup β pick by your device
*Wikipedia-Korean perplexity, lower is better. Q4_K_M = 5.79 baseline. English builds are tuned on English; see each repo. β Separate 80-chunk run (40,960 tokens) on a different model β compare within a model, not across rows.
π Why MLX for Korean but GGUF for English on iPhone? Apple-native MLX only does uniform quantization. Korean survives it; English needs our proprietary quantization, which only GGUF supports β so the English iPhone build ships as a GGUF you run with PocketPal. Honest, not lazy.
π POCKET-26B β a Gemma4-26B-A4B-based sibling that loads in any app today (Ollama Β· LM Studio Β· PocketPal Β· MLX), no bleeding-edge runtime needed: [GGUF](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF) (Q2_K11 GB Β·Q4_K_M17 GB Β· GPQA-Diamond 67%). Universal compatibility for 12 GB phones, PC, and browser.
Benchmarks β what is measured, what is not
We measure Bonsai on the same machine with the same stock `llama.cpp`, and we tell you where we lose.
[measured] Generation speed β POCKET wins on both CPU and GPU:
[measured on a MacBook M3 Pro, 18 GB] β and on a laptop, POCKET wins every axis, including prompt processing:
On a laptop GPU the arithmetic headroom that let Bonsai win prefill on an H100 is gone, so MoE sparsity wins across the board. POCKET-35B-Q2_K runs on the M3 Pro's CPU at 19.5 tok/s β on an 18 GB Mac, run Q2_K on CPU (-ngl 0); its 13 GB exceeds the recommended Metal budget.
[measured β GPQA Diamond, 198q, greedy] reasoning quality vs quantization:
[pending β community reports welcome] on-device iPhone and Strix Halo throughput. We publish only what we ran ourselves; help us fill the rest.
The same-size rival Ternary-Bonsai-27B-Q2_0 (7.2 GB) fails to load in upstream llama.cpp β it needs the PrismML fork. POCKET runs on the tools you already have.Files in this repo
Quickstart β no fork needed
# any recent llama.cpp β brew / winget / apt, or LM Studio / Ollama
llama-cli -m POCKET-35B-Q2_K.gguf -p "μλ
νμΈμ" -ngl 0 -t 8
# reproduce our CPU numbers:
llama-bench -m POCKET-35B-IQ1_M.gguf -p 128 -n 64 -ngl 0 -t 16Use physical-core count for -t (max ~32). Do not pass all threads β it can slow down sharply.
Lineage β where POCKET comes from
POCKET is quantized from [Darwin-36B-Opus](https://huggingface.co/FINAL-Bench/Darwin-36B-Opus), VIDRAFT's flagship β a model bred and evolved over several generations on the Darwin platform (crossbreeding, healing, expert surgery). Darwin-36B-Opus itself traces back to a Qwen3.5-family MoE architecture.
The CPU/GPU speed comes from the sparse-MoE architecture plus ordinary quantization β reproducible with the same base and the same tools. What we add is the Darwin-evolved weights, the honest measurement, the Korean tuning, and the pruning that makes the 5 GB phone builds.
Limitations
- The iPhone/Mac speed is not yet measured by us β community reports welcome.
- Extreme quants (
IQ1_M) hurt Korean ~2.8Γ more than English; useQ2_Kor larger for quality. - English phone builds trade quality for size; the PC build (
PC-mix) is much closer to full quality.
License
Apache-2.0.
POCKET is a VIDRAFT model family. 35B, in your pocket. No GPU.
Learn more
- On-device LLMs without a GPU β and how POCKET measures up: Can you run a large LLM without a GPU?
- What model quantization is, and why a 4-bit model stays smart: What is model quantization?
<!-- POCKET-FAMILY -->
π§© The POCKET Family β On-device AI by VIDRAFT
Big models, small hardware. No GPU, no cloud.
Models
- π¦ POCKET-35B-GGUF β flagship, PC / server, no GPU
- π¦ POCKET-26B-GGUF β compact 26B
- π°π· POCKET-KR-GGUF β Korean, Android
- π POCKET-KR-MLX β Korean, iPhone / Mac
- π POCKET-EN-GGUF β English, phone / PC
- π» POCKET-Qwen3.8-Flash-Next-GGUF β 180B on a laptop (8 GB VRAM + 32 GB RAM)
- πΌοΈ POCKET-Image-Zimage β character-perfect text in any image
- π₯οΈ POCKET-Zimage-CPU β photoreal images on a CPU only
Demos & tools (Spaces)
- π¨ POCKET-Image Studio β text-in-image, generate in-page
- π₯οΈ POCKET-35B-CPU β 35B answering on a CPU
- π₯οΈ POCKET-26B-CPU β 26B on a CPU
- πΌοΈ POCKET-Zimage-CPU β image generation on a CPU
<!-- /POCKET-FAMILY -->
