Vrushali777/vllm-inference-benchmark
vLLM Inference Benchmarking
A hands-on benchmarking project comparing vLLM against HuggingFace Transformers across two core innovations: PagedAttention and Continuous Batching. Results are visualised in an interactive Gradio dashboard.
Project Structure
vLLM_Benchmarking/
├── app.py # Gradio dashboard
├── data/
│ ├── paged_attention_results.csv
│ └── batching_results.csv
├── experiments.ipynb # Benchmark notebooks
└── README.mdHardware
Experiments
1. PagedAttention — vLLM vs HuggingFace Transformers
Goal: Measure whether vLLM's PagedAttention algorithm delivers throughput and latency improvements over a naive HuggingFace Transformers baseline across different model sizes.
Models tested:
Qwen1.5-0.5BQwen2.5-1.5BQwen2.5-7B
Batch sizes: 10, 50, 100
Metrics collected:
time_sec— wall-clock inference time for the batchthroughput_tokens_per_sec— total tokens generated per secondgpu_mem_used_mb— GPU memory reported by PyTorch allocator
Key finding: vLLM's advantage scales with model size. On Qwen1.5-0.5B and Qwen2.5-1.5B, the KV-cache is small enough that memory fragmentation is not a bottleneck, and HuggingFace is competitive or faster. On Qwen2.5-7B, vLLM achieves up to 11.94× higher throughput at batch size 50 — the point at which memory pressure makes PagedAttention's non-contiguous KV-cache management decisive.
Note on memory metrics: Dynamic GPU memory fragmentation was not captured during inference. The gpu_mem_used_mb column reflects PyTorch's allocator reporting, which shows near-zero for HuggingFace (a known allocator quirk) and 0 for vLLM. Throughput and latency serve as proxy metrics for PagedAttention's memory efficiency — vLLM's superior throughput at larger batch sizes is the observable evidence of its ability to fit more concurrent requests into the same VRAM budget.
2. Batching Strategies
Goal: Compare three request scheduling strategies to understand the throughput impact of continuous batching vs static batching.
Strategies:
single_request— requests processed one at a time, sequentiallystatic_batch— requests grouped into a fixed batch; GPU waits for all sequences to finish before accepting new requestscontinuous_batching— vLLM's approach; as soon as any sequence in the batch finishes, its GPU slot is immediately filled with the next queued request
Batch sizes: 1, 10, 50
Metrics collected:
time_sec— wall-clock time for the batch windowthroughput_tokens_per_sec— total tokens generated per second
Key finding: Continuous batching delivers 38–39× higher throughput vs static batching at batch sizes 1 and 10, and 22.5× at batch size 50.
Note on latency: Wall-clock time_sec is not a fair per-request latency comparison for continuous batching. Because continuous batching keeps filling GPU slots as requests complete, it processes more than the nominal batch size within each time window — meaning it is doing more total work. Throughput (tokens/sec) is the appropriate metric for evaluating continuous batching performance.
Dashboard
The Gradio dashboard (app.py) is a static results viewer — no inference is performed at runtime. It reads from the two CSV files in data/ and renders interactive Plotly charts.
Tabs:
- PagedAttention — model dropdown (Qwen1.5-0.5B / Qwen2.5-1.5B / Qwen2.5-7B), throughput comparison, latency comparison, and speedup ratio chart. Defaults to Qwen2.5-7B.
- Batching Strategies — throughput by strategy, and continuous vs static speedup ratio.
- Concepts — written explanation of PagedAttention and continuous batching, tied back to benchmark findings.
To run:
pip install gradio pandas plotly
python app.pyResults Summary
🔗 Interactive Dashboard: 👉 https://huggingface.co/spaces/Vrushali777/vllm-inference-benchmark
PagedAttention — vLLM vs HuggingFace
Throughput Comparison (Qwen2.5-7B)
Latency Comparison (Qwen2.5-7B)
Speedup (vLLM vs HF) (Qwen2.5-7B)
Batching — Continuous vs Static Throughput Speedup
Throughput by Strategy
Continuous vs Static Speedup
Key Takeaways
- PagedAttention's benefit is model-size dependent. It solves a memory fragmentation problem that only becomes a bottleneck when the KV-cache is large — i.e., large models at high concurrency. Benchmarking small models alongside large ones is intentional: it shows the threshold at which the technique becomes impactful.
- Continuous batching's advantage is throughput, not individual request latency. The metric that matters for production serving is how many total requests the system can handle per second — and continuous batching wins decisively on this across all batch sizes tested.
- vLLM's OpenAI-compatible API means these gains are accessible as a drop-in replacement for any system already built against the OpenAI SDK.
References
- vLLM: Efficient Memory Management for LLM Serving with PagedAttention — Kwon et al., UC Berkeley, 2023
- vLLM Documentation
- vLLM GitHub
- vLLM Red Hat Blog
