Team Ai
Modelpublic

diffuse-cpp/LLaDA-8B-Instruct-GGUF

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
7likes238downloads
README.md101 linesDownload Raw Back to root
1---2license: apache-2.03tags:4  - diffusion5  - llada6  - gguf7  - cpu-inference8  - diffuse-cpp9language:10  - en11base_model: GSAI-ML/LLaDA-8B-Instruct12pipeline_tag: text-generation13---14 15# LLaDA-8B-Instruct-GGUF16 17GGUF quantizations of [GSAI-ML/LLaDA-8B-Instruct](https://huggingface.co/GSAI-ML/LLaDA-8B-Instruct) for use with [diffuse-cpp](https://github.com/iafiscal1212/diffuse-cpp), the first C++ inference engine for Diffusion Language Models.18 19LLaDA is a masked diffusion language model based on the Llama backbone. Unlike autoregressive models that generate one token at a time, LLaDA generates all tokens in parallel through iterative refinement — making it compute-bound rather than memory-bound on CPU.20 21**On a 12-core CPU, LLaDA with diffuse-cpp reaches 27.7 tok/s on translation tasks — 3.3x faster than llama.cpp (8.51 tok/s) on the same hardware.**22 23## Available Quantizations24 25| File | Type | Size | Description |26|------|------|------|-------------|27| `llada-8b-f16.gguf` | F16 | ~14.9 GB | Full precision, best quality |28| `llada-8b-q8_0.gguf` | Q8_0 | ~8.4 GB | 8-bit quantization, near-lossless |29| `llada-8b-q4km.gguf` | Q4_K_M | ~5.1 GB | 4-bit mixed, best speed/quality ratio |30 31**Recommended:** Q4_K_M for most users.32 33## Quick Start34 35```bash36# Download37huggingface-cli download diffuse-cpp/LLaDA-8B-Instruct-GGUF llada-8b-q4km.gguf38 39# Build diffuse-cpp40git clone --recursive https://github.com/iafiscal1212/diffuse-cpp.git41cd diffuse-cpp42cmake -B build -DCMAKE_BUILD_TYPE=Release43cmake --build build -j$(nproc)44 45# Run46./build/diffuse-cli -m ../llada-8b-q4km.gguf \47    --tokens "128000,3923,374,279,6864,315,9822,30" \48    -n 256 -s 16 -t 12 --remasking entropy_exit49```50 51## Performance52 53Benchmarked on AMD EPYC 4465P 12-Core, Q4_K_M, entropy_exit + inter-step cache, B=256:54 55| Prompt | No-Cache | Cache | Steps | vs llama.cpp |56|--------|----------|-------|-------|-------------|57| Capital of France? | 17.5 | **24.4 tok/s** | 3 | 2.9x |58| Translate to French | 25.9 | **27.7 tok/s** | 2 | **3.3x** |59| 15 x 23? | 12.8 | **15.7 tok/s** | 4 | 1.8x |60| Translate to Spanish | 7.6 | **22.9 tok/s** | 7 | 2.7x |61| Python is_prime() | 3.2 | **4.9 tok/s** | 16 | 0.6x |62| Poem about ocean | 3.2 | **5.3 tok/s** | 16 | 0.6x |63| Why is sky blue? | 3.3 | **12.0 tok/s** | 16 | 1.4x |64| List the planets | 3.3 | **9.4 tok/s** | 15 | 1.1x |65| **Average** | **9.6** | **15.3 tok/s** | | **1.8x** |66 67- Inter-step cache: 1.6x average speedup with no quality degradation68- 6 of 8 prompts outperform llama.cpp (8.51 tok/s baseline)69- LLaDA excels at translation tasks (converges in 2-5 steps)70 71## Model Details72 73- **Architecture:** Llama backbone with bidirectional (non-causal) attention74- **Parameters:** 8B75- **Layers:** 3276- **Hidden size:** 409677- **Attention:** MHA (32 query heads, 32 KV heads)78- **FFN:** SwiGLU, intermediate 1228879- **Vocabulary:** 126,464 tokens80- **RoPE theta:** 500,00081- **Mask token ID:** 12633682 83## Also Available84 85- **[Dream-v0-Instruct-7B-GGUF](https://huggingface.co/diffuse-cpp/Dream-v0-Instruct-7B-GGUF)** — Qwen2.5 backbone, GQA. Excels at math and code (21.6 tok/s, correctly solves arithmetic in 2 steps).86 87## Citation88 89```bibtex90@software{diffuse_cpp_2026,91  title={diffuse-cpp: High-Performance Inference for Diffusion Language Models},92  author={Carmen Esteban},93  year={2026},94  url={https://github.com/iafiscal1212/diffuse-cpp}95}96```97 98## License99 100Apache 2.0101