Team Ai
Datasetpublic

dispatchAI/on-device-latency

On-Device Latency Benchmark Real-world inference latency data for mobile-optimized LLMs, measured on actual phone hardware. Hardware Spec Value Device Samsung S20 FE 5G SoC Snapdragon 865 RAM 8GB OS Android 13 Runtime llama.cpp (4 threads) Metrics tokens_per_sec — Generation speed during inference latency_ms_per_token — Time per generated token ram_usage_mb — Peak RAM during inference file_size_mb — GGUF model file size… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/on-device-latency.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes23downloads

dispatchAI/on-device-latency · main · files are served by the source, never re-hosted here