dispatchAI/Llama-3.2-1B-Instruct-Q4-mobile
083
Llama 3.2 1B Instruct - Q4 Mobile (GGUF)
Meta's Llama 3.2 1B Instruct, quantized to INT4 GGUF format for mobile deployment by Dispatch AI.
Why This Model?
Mobile-optimized for deployment on Android phones (Snapdragon 865+), laptops, IoT devices, and any hardware with 4GB+ RAM. No GPU required.
Performance on Samsung S20 FE (Snapdragon 865)
Use Cases
- Chatbots & conversational AI on mobile devices
- Instruction following in resource-constrained environments
- Content summarization, text classification, RAG pipelines
- Educational apps, tutoring systems
Quick Start
# Install llama.cpp
git clone https://github.com/ggerganov/llama.cpp && cd llama.cpp && cmake -B build -DLLAMA_NATIVE=ON && cmake --build build --config Release
# Download this model
huggingface-cli download dispatchAI/Llama-3.2-1B-Instruct-Q4-mobile ggml-model-Q4_K_M.gguf --local-dir ./models
# Run inference immediately
./build/bin/main -m ./models/ggml-model-Q4_K_M.gguf -p "Hello" -n 256 -t 4Hardware Requirements
Limitations
- ~5% quality degradation vs FP16 on complex reasoning tasks
- Not suitable for high-precision numerical computation
- Context window follows base model (~128K tokens)
About Dispatch AI
Re-engineering LLMs for mobile and edge deployment. HuggingFace - 40+ models, 13K+ downloads
