Team Ai
Modelpublic

FINAL-Bench/POCKET-35B-GGUF

sourceHugging Faceapache-2.0updated 15d agoView on Hugging Face
84likes554kdownloads
Model Card
### πŸ†• [POCKET-Qwen3.8-Flash-Next](https://huggingface.co/FINAL-Bench/POCKET-Qwen3.8-Flash-Next-GGUF) β€” a 180B model running on a laptop with 8 GB VRAM + 32 GB RAM Β· 4.17 tok/s measured. ![New](https://huggingface.co/FINAL-Bench/POCKET-Qwen3.8-Flash-Next-GGUF) ![VRAM]() ![RAM]() ![Speed]() <!-- POCKET-FLASHNEXT-BADGE -->
### πŸ†• [POCKET-Zimage-CPU](https://huggingface.co/FINAL-Bench/POCKET-Zimage-CPU) β€” photoreal images in 46 s on a CPU only. No GPU, no CUDA, no Python. ![New](https://huggingface.co/FINAL-Bench/POCKET-Zimage-CPU) ![Space](https://huggingface.co/spaces/FINAL-Bench/POCKET-Zimage-CPU) ![RAM]() <!-- POCKET-Zimage-CPU-BADGE -->
### πŸ“š Collections β–Ά [POCKET Models](https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6) β€” this family (on-device, no GPU) Darwin Family Β· Aether Foundation Β· VKAE Accelerated

[image]

POCKET-35B-GGUF

A 35B model that runs on your PC with no GPU β€” and on your phone. Just stock llama.cpp. No fork, no CUDA, no cloud.

πŸš€ Try it live, no install β†’ ![POCKET-35B demo](https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU) ![POCKET-26B demo](https://huggingface.co/spaces/FINAL-Bench/POCKET-26B-CPU) β€” both answering on a CPU-only box (no GPU). POCKET-26B is Gemma4-based.

![License](https://www.apache.org/licenses/LICENSE-2.0) ![Runtime](https://github.com/ggml-org/llama.cpp) ![No GPU]() ![Base]()

Pick your build β†’ ![35B](https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF) ![26B](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF) ![KR GGUF](https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF) ![KR MLX-0f6e56)](https://huggingface.co/FINAL-Bench/POCKET-KR-MLX) ![EN GGUF](https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF) ![180B laptop](https://huggingface.co/FINAL-Bench/POCKET-Qwen3.8-Flash-Next-GGUF) ![Image NF4](https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage) ![Image CPU](https://huggingface.co/FINAL-Bench/POCKET-Zimage-CPU)

The POCKET lineup β€” pick by your device

RepoFileSizeRuns onBest forKorean PPL*
POCKET-35B-GGUFQ4_K_M21 GBPC / server (32 GB RAM)top quality5.79
POCKET-35B-GGUFQ2_K ⭐13 GBmini-PC, no GPUdaily driver6.49
POCKET-35B-GGUFIQ1_M8.2 GB16 GB RAM boxsmallest full model9.69
POCKET-KR-GGUFIQ2_M5.1 GBAndroid 8 GB+πŸ‡°πŸ‡· Korean phone7.95
POCKET-KR-MLX2-bit5.1 GB🍎 iPhone / iPad / MacπŸ‡°πŸ‡· Korean, Apple-native7.95
POCKET-EN-GGUFiPhone-mix5.3 GB🍎 iPhone (PocketPal)🌍 English phoneβ€”
POCKET-EN-GGUFPC-mix6.8 GBPC / Android🌍 English, best qualityβ€”
POCKET-Qwen3.8-Flash-Next-GGUFQ4_K_M111 GiBπŸ’» laptop, 8 GB VRAM + 32 GB RAM180B on a laptop6.03†

*Wikipedia-Korean perplexity, lower is better. Q4_K_M = 5.79 baseline. English builds are tuned on English; see each repo. †Separate 80-chunk run (40,960 tokens) on a different model β€” compare within a model, not across rows.

🍎 Why MLX for Korean but GGUF for English on iPhone? Apple-native MLX only does uniform quantization. Korean survives it; English needs our proprietary quantization, which only GGUF supports β€” so the English iPhone build ships as a GGUF you run with PocketPal. Honest, not lazy.
πŸ†• POCKET-26B β€” a Gemma4-26B-A4B-based sibling that loads in any app today (Ollama Β· LM Studio Β· PocketPal Β· MLX), no bleeding-edge runtime needed: [GGUF](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF) (Q2_K 11 GB Β· Q4_K_M 17 GB Β· GPQA-Diamond 67%). Universal compatibility for 12 GB phones, PC, and browser.

[image]

Benchmarks β€” what is measured, what is not

We measure Bonsai on the same machine with the same stock `llama.cpp`, and we tell you where we lose.

[measured] Generation speed β€” POCKET wins on both CPU and GPU:

POCKET-35B IQ1_MBonsai-27B Q1_0
CPU generate (Xeon, 16t)27.0 tok/s10.1🟒 2.69Γ—
GPU generate (H100)197 tok/s89🟒 2.22Γ—
GPU prompt (H100)7531816πŸ”΄ 0.41Γ—
Quality (HellaSwag, 400q)61.0%60.0%βšͺ tie (CI overlaps)

[measured on a MacBook M3 Pro, 18 GB] β€” and on a laptop, POCKET wins every axis, including prompt processing:

POCKET-35B IQ1_MBonsai-27B Q1_0
Metal generate (tg64)25.4 tok/s12.8🟒 1.99Γ—
CPU generate (8 threads)13.8 tok/s4.4🟒 3.13Γ—
Metal prompt (pp128)240.7 tok/s73.4🟒 3.28Γ—
CPU prompt (pp128)45.5 tok/s9.6🟒 4.75Γ—

On a laptop GPU the arithmetic headroom that let Bonsai win prefill on an H100 is gone, so MoE sparsity wins across the board. POCKET-35B-Q2_K runs on the M3 Pro's CPU at 19.5 tok/s β€” on an 18 GB Mac, run Q2_K on CPU (-ngl 0); its 13 GB exceeds the recommended Metal budget.

[measured β€” GPQA Diamond, 198q, greedy] reasoning quality vs quantization:

ModelGPQA-Diamond (greedy)
Qwen3.6-35B-A3B73.2%
POCKET-35B Q4KM68.7%
POCKET-35B Q2_K60.1%

[pending β€” community reports welcome] on-device iPhone and Strix Halo throughput. We publish only what we ran ourselves; help us fill the rest.

The same-size rival Ternary-Bonsai-27B-Q2_0 (7.2 GB) fails to load in upstream llama.cpp β€” it needs the PrismML fork. POCKET runs on the tools you already have.

Files in this repo

FileSizebpwRuns onKorean PPL
POCKET-35B-Q4_K_M.gguf21 GB4.5PC 32 GB RAM5.79 (top)
POCKET-35B-Q3_K_M.gguf16 GB3.4PC 24 GB6.06
`POCKET-35B-Q2_K.gguf` ⭐13 GB2.6mini-PC 16–24 GB6.49 (best value)
POCKET-35B-IQ1_M.gguf8.2 GB1.916 GB RAM9.69 (smallest)

Quickstart β€” no fork needed

bash
# any recent llama.cpp β€” brew / winget / apt, or LM Studio / Ollama
llama-cli -m POCKET-35B-Q2_K.gguf -p "μ•ˆλ…•ν•˜μ„Έμš”" -ngl 0 -t 8
# reproduce our CPU numbers:
llama-bench -m POCKET-35B-IQ1_M.gguf -p 128 -n 64 -ngl 0 -t 16

Use physical-core count for -t (max ~32). Do not pass all threads β€” it can slow down sharply.

Lineage β€” where POCKET comes from

POCKET is quantized from [Darwin-36B-Opus](https://huggingface.co/FINAL-Bench/Darwin-36B-Opus), VIDRAFT's flagship β€” a model bred and evolved over several generations on the Darwin platform (crossbreeding, healing, expert surgery). Darwin-36B-Opus itself traces back to a Qwen3.5-family MoE architecture.

ComponentOrigin
Starting checkpoint[Darwin-36B-Opus](https://huggingface.co/FINAL-Bench/Darwin-36B-Opus) β€” VIDRAFT, multi-generation Darwin evolution
Base architectureQwen3.5-family MoE (256 experts, top-8), unchanged
Quantization (Q4_K_M…IQ1_M)stock llama.cpp β€” no custom format
Runtimeupstream llama.cpp / Apple MLX β€” unmodified
Proprietary language-specific tuning (KR/EN builds)ours (VIDRAFT)

The CPU/GPU speed comes from the sparse-MoE architecture plus ordinary quantization β€” reproducible with the same base and the same tools. What we add is the Darwin-evolved weights, the honest measurement, the Korean tuning, and the pruning that makes the 5 GB phone builds.

Limitations

  • β€”The iPhone/Mac speed is not yet measured by us β€” community reports welcome.
  • β€”Extreme quants (IQ1_M) hurt Korean ~2.8Γ— more than English; use Q2_K or larger for quality.
  • β€”English phone builds trade quality for size; the PC build (PC-mix) is much closer to full quality.

License

Apache-2.0.


POCKET is a VIDRAFT model family. 35B, in your pocket. No GPU.

Learn more

<!-- POCKET-FAMILY -->


🧩 The POCKET Family β€” On-device AI by VIDRAFT

Big models, small hardware. No GPU, no cloud.

Models

Demos & tools (Spaces)

πŸ“š Full POCKET collection

<!-- /POCKET-FAMILY -->