Team Ai
Apppublic

Vrushali777/vllm-inference-benchmark

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

vLLM Inference Benchmarking

A hands-on benchmarking project comparing vLLM against HuggingFace Transformers across two core innovations: PagedAttention and Continuous Batching. Results are visualised in an interactive Gradio dashboard.


Project Structure

vLLM_Benchmarking/
├── app.py                          # Gradio dashboard
├── data/
│   ├── paged_attention_results.csv
│   └── batching_results.csv
├── experiments.ipynb               # Benchmark notebooks
└── README.md

Hardware

SpecValue
GPUNVIDIA A100-SXM4-80GB
VRAM80 GB
CUDA13.0
Driver580.82.07
TDP400 W

Experiments

1. PagedAttention — vLLM vs HuggingFace Transformers

Goal: Measure whether vLLM's PagedAttention algorithm delivers throughput and latency improvements over a naive HuggingFace Transformers baseline across different model sizes.

Models tested:

  • —Qwen1.5-0.5B
  • —Qwen2.5-1.5B
  • —Qwen2.5-7B

Batch sizes: 10, 50, 100

Metrics collected:

  • —time_sec — wall-clock inference time for the batch
  • —throughput_tokens_per_sec — total tokens generated per second
  • —gpu_mem_used_mb — GPU memory reported by PyTorch allocator

Key finding: vLLM's advantage scales with model size. On Qwen1.5-0.5B and Qwen2.5-1.5B, the KV-cache is small enough that memory fragmentation is not a bottleneck, and HuggingFace is competitive or faster. On Qwen2.5-7B, vLLM achieves up to 11.94× higher throughput at batch size 50 — the point at which memory pressure makes PagedAttention's non-contiguous KV-cache management decisive.

Note on memory metrics: Dynamic GPU memory fragmentation was not captured during inference. The gpu_mem_used_mb column reflects PyTorch's allocator reporting, which shows near-zero for HuggingFace (a known allocator quirk) and 0 for vLLM. Throughput and latency serve as proxy metrics for PagedAttention's memory efficiency — vLLM's superior throughput at larger batch sizes is the observable evidence of its ability to fit more concurrent requests into the same VRAM budget.


2. Batching Strategies

Goal: Compare three request scheduling strategies to understand the throughput impact of continuous batching vs static batching.

Strategies:

  • —single_request — requests processed one at a time, sequentially
  • —static_batch — requests grouped into a fixed batch; GPU waits for all sequences to finish before accepting new requests
  • —continuous_batching — vLLM's approach; as soon as any sequence in the batch finishes, its GPU slot is immediately filled with the next queued request

Batch sizes: 1, 10, 50

Metrics collected:

  • —time_sec — wall-clock time for the batch window
  • —throughput_tokens_per_sec — total tokens generated per second

Key finding: Continuous batching delivers 38–39× higher throughput vs static batching at batch sizes 1 and 10, and 22.5× at batch size 50.

Note on latency: Wall-clock time_sec is not a fair per-request latency comparison for continuous batching. Because continuous batching keeps filling GPU slots as requests complete, it processes more than the nominal batch size within each time window — meaning it is doing more total work. Throughput (tokens/sec) is the appropriate metric for evaluating continuous batching performance.


Dashboard

The Gradio dashboard (app.py) is a static results viewer — no inference is performed at runtime. It reads from the two CSV files in data/ and renders interactive Plotly charts.

Tabs:

  • —PagedAttention — model dropdown (Qwen1.5-0.5B / Qwen2.5-1.5B / Qwen2.5-7B), throughput comparison, latency comparison, and speedup ratio chart. Defaults to Qwen2.5-7B.
  • —Batching Strategies — throughput by strategy, and continuous vs static speedup ratio.
  • —Concepts — written explanation of PagedAttention and continuous batching, tied back to benchmark findings.

To run:

bash
pip install gradio pandas plotly
python app.py

Results Summary

🔗 Interactive Dashboard: 👉 https://huggingface.co/spaces/Vrushali777/vllm-inference-benchmark


PagedAttention — vLLM vs HuggingFace

Throughput Comparison (Qwen2.5-7B)

[image]

Latency Comparison (Qwen2.5-7B)

[image]

Speedup (vLLM vs HF) (Qwen2.5-7B)

[image]

ModelBest SpeedupAt Batch Size
Qwen1.5-0.5B0.08×— (HF faster)
Qwen2.5-1.5B0.56×— (HF faster)
Qwen2.5-7B11.94×50

Batching — Continuous vs Static Throughput Speedup

Throughput by Strategy

[image]

Continuous vs Static Speedup

[image]

Batch SizeSpeedup
138.2×
1039.1×
5022.5×

Key Takeaways

  • —PagedAttention's benefit is model-size dependent. It solves a memory fragmentation problem that only becomes a bottleneck when the KV-cache is large — i.e., large models at high concurrency. Benchmarking small models alongside large ones is intentional: it shows the threshold at which the technique becomes impactful.
  • —Continuous batching's advantage is throughput, not individual request latency. The metric that matters for production serving is how many total requests the system can handle per second — and continuous batching wins decisively on this across all batch sizes tested.
  • —vLLM's OpenAI-compatible API means these gains are accessible as a drop-in replacement for any system already built against the OpenAI SDK.

References