ninjals/llm-inference-benchmark-playground
0
LLM Inference & Benchmark Playground
A lightweight Hugging Face ZeroGPU demo Space for interactive benchmarking of remote / OpenAI-compatible LLM endpoints.
This Space is the demo deployment of the broader project. It focuses on:
- single-prompt benchmarking
- batch / multi-prompt benchmarking
- remote batch-size sweep
- benchmark history, plots, and export workflows
For the full local-vLLM version with richer backend routing, local serving workflows, and deeper telemetry, use the Docker / paid-GPU deployment.
ZeroGPU demo scope
This Space is designed to run on Gradio + ZeroGPU without loading local CUDA engines such as vLLM at startup.
Included in this ZeroGPU demo
- Remote / OpenAI-compatible API benchmarking
- Single-prompt benchmarking
- Batch / multi-prompt benchmarking
- Remote batch-size sweep
- Batch prompt input supporting:
- one prompt per line
###-delimited multi-line prompts- Latency / p50 latency / output-token / tokens-per-second metrics
- Benchmark history tracking
- CSV export
- JSON export
- Markdown benchmark report export
- Session bundle export (
.zip) - Simple provider-level throughput comparison plots
Not included in this ZeroGPU demo
- Local vLLM engine loading
- Local GPU telemetry
- TTFT tied to local serving internals
- Full local backend reload / engine management workflows
- The richer local-serving path from the Docker version
Recommended use
Use this Space when you want to:
- quickly demonstrate the benchmark UI
- test remote OpenAI-compatible endpoints
- showcase benchmark / export / report workflows
- keep a lightweight protected or public demo online
Full project structure
This project currently has two deployment paths:
1. ZeroGPU Gradio demo Space
Best for:
- easy online demo
- protected/public showcase
- remote API benchmarking
- light deployment footprint
2. Full Docker / paid-GPU deployment
Best for:
- local vLLM execution
- local GPU telemetry
- fuller backend-routing workflows
- richer serving-side benchmarking
Supported API presets in the demo
Depending on what endpoint you provide, the ZeroGPU demo can benchmark:
- OpenAI API
- DeepSeek API
- LM Studio OpenAI bridge
- vLLM OpenAI server
- Ollama OpenAI bridge
- Custom OpenAI-compatible endpoints
Benchmark input flexibility
Batch prompts support two formats.
One prompt per line
Use this for short prompts.
###-delimited blocks
Use this for longer multi-line prompts.
Example:
###
Explain why KV cache matters for transformer inference.
What deployment problem does it solve?
###
What is the difference between latency and throughput in LLM serving?
How should engineers balance them?
###
Why does batch size affect GPU utilization?
How should batch size be selected in practice?