dispatchAI/on-device-latency
On-Device Latency Benchmark Real-world inference latency data for mobile-optimized LLMs, measured on actual phone hardware. Hardware Spec Value Device Samsung S20 FE 5G SoC Snapdragon 865 RAM 8GB OS Android 13 Runtime llama.cpp (4 threads) Metrics tokens_per_sec — Generation speed during inference latency_ms_per_token — Time per generated token ram_usage_mb — Peak RAM during inference file_size_mb — GGUF model file size… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/on-device-latency.
On-Device Latency Benchmark
Real-world inference latency data for mobile-optimized LLMs, measured on actual phone hardware.
Hardware
Metrics
- tokens_per_sec — Generation speed during inference
- latency_ms_per_token — Time per generated token
- ram_usage_mb — Peak RAM during inference
- file_size_mb — GGUF model file size
Usage
from datasets import load_dataset
ds = load_dataset("dispatchAI/on-device-latency")Note
These are estimated benchmarks based on model architecture, parameter count, and quantization level. Actual on-device measurements may vary. Real phone farm benchmarks are ongoing.
🚀 dispatchAI — Small. Mobile. Free. UAE-built.
