Team Ai
Apppublic

ninjals/llm-inference-benchmark-playground

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes
App README

LLM Inference & Benchmark Playground

A lightweight Hugging Face ZeroGPU demo Space for interactive benchmarking of remote / OpenAI-compatible LLM endpoints.

This Space is the demo deployment of the broader project. It focuses on:

  • —single-prompt benchmarking
  • —batch / multi-prompt benchmarking
  • —remote batch-size sweep
  • —benchmark history, plots, and export workflows

For the full local-vLLM version with richer backend routing, local serving workflows, and deeper telemetry, use the Docker / paid-GPU deployment.

ZeroGPU demo scope

This Space is designed to run on Gradio + ZeroGPU without loading local CUDA engines such as vLLM at startup.

Included in this ZeroGPU demo

  • —Remote / OpenAI-compatible API benchmarking
  • —Single-prompt benchmarking
  • —Batch / multi-prompt benchmarking
  • —Remote batch-size sweep
  • —Batch prompt input supporting:
  • —one prompt per line
  • —###-delimited multi-line prompts
  • —Latency / p50 latency / output-token / tokens-per-second metrics
  • —Benchmark history tracking
  • —CSV export
  • —JSON export
  • —Markdown benchmark report export
  • —Session bundle export (.zip)
  • —Simple provider-level throughput comparison plots

Not included in this ZeroGPU demo

  • —Local vLLM engine loading
  • —Local GPU telemetry
  • —TTFT tied to local serving internals
  • —Full local backend reload / engine management workflows
  • —The richer local-serving path from the Docker version

Recommended use

Use this Space when you want to:

  • —quickly demonstrate the benchmark UI
  • —test remote OpenAI-compatible endpoints
  • —showcase benchmark / export / report workflows
  • —keep a lightweight protected or public demo online

Full project structure

This project currently has two deployment paths:

1. ZeroGPU Gradio demo Space

Best for:

  • —easy online demo
  • —protected/public showcase
  • —remote API benchmarking
  • —light deployment footprint

2. Full Docker / paid-GPU deployment

Best for:

  • —local vLLM execution
  • —local GPU telemetry
  • —fuller backend-routing workflows
  • —richer serving-side benchmarking

Supported API presets in the demo

Depending on what endpoint you provide, the ZeroGPU demo can benchmark:

  • —OpenAI API
  • —DeepSeek API
  • —LM Studio OpenAI bridge
  • —vLLM OpenAI server
  • —Ollama OpenAI bridge
  • —Custom OpenAI-compatible endpoints

Benchmark input flexibility

Batch prompts support two formats.

One prompt per line

Use this for short prompts.

###-delimited blocks

Use this for longer multi-line prompts.

Example:

text
###
Explain why KV cache matters for transformer inference.
What deployment problem does it solve?
###
What is the difference between latency and throughput in LLM serving?
How should engineers balance them?
###
Why does batch size affect GPU utilization?
How should batch size be selected in practice?