Team Ai
Modelpublic

dispatchAI/Llama-3.2-1B-Instruct-Q4-mobile

sourceHugging Facellama3.2updated 3mo agoView on Hugging Face
0likes83downloads
Model Card

Llama 3.2 1B Instruct - Q4 Mobile (GGUF)

Meta's Llama 3.2 1B Instruct, quantized to INT4 GGUF format for mobile deployment by Dispatch AI.

PropertyValue
Basemeta-llama/Llama-3.2-1B-Instruct
Parameters1.23 billion
QuantizationQ4KM (4-bit k-means)
Size~767 MB
FormatGGUF (llama.cpp)
LicenseLlama 3.2 Community

Why This Model?

Mobile-optimized for deployment on Android phones (Snapdragon 865+), laptops, IoT devices, and any hardware with 4GB+ RAM. No GPU required.

Performance on Samsung S20 FE (Snapdragon 865)

MetricThis VersionOriginal FP16
Size767 MB~2.5 GB
Speed~28 tok/s CPU~8 tok/s
Memory~1.2 GB~3.8 GB
Quality~95% of original100% baseline

Use Cases

  • —Chatbots & conversational AI on mobile devices
  • —Instruction following in resource-constrained environments
  • —Content summarization, text classification, RAG pipelines
  • —Educational apps, tutoring systems

Quick Start

bash
# Install llama.cpp
git clone https://github.com/ggerganov/llama.cpp && cd llama.cpp && cmake -B build -DLLAMA_NATIVE=ON && cmake --build build --config Release

# Download this model
huggingface-cli download dispatchAI/Llama-3.2-1B-Instruct-Q4-mobile ggml-model-Q4_K_M.gguf --local-dir ./models

# Run inference immediately
./build/bin/main -m ./models/ggml-model-Q4_K_M.gguf -p "Hello" -n 256 -t 4

Hardware Requirements

RequirementMinimumRecommended
RAM4 GB6 GB+
Storage1 GB free2 GB+
CPU4-core ARM64/x86_648-core Snapdragon 865+
GPUNot requiredAny (faster)

Limitations

  • —~5% quality degradation vs FP16 on complex reasoning tasks
  • —Not suitable for high-precision numerical computation
  • —Context window follows base model (~128K tokens)

About Dispatch AI

Re-engineering LLMs for mobile and edge deployment. HuggingFace - 40+ models, 13K+ downloads