Team Ai
Apppublic

Kotgire58/InferenceBenchmark

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes
App README

๐Ÿš€ Inference Benchmarking Suite (Streamlit Edition)

A modular, educational benchmarking toolkit designed to teach and demonstrate modern LLM inference optimizations โ€” including batching, KV-cache reuse, speculative decoding, vLLM acceleration, and real-time chatbot comparisons โ€” with a clean Streamlit UI.

This project is structured for hands-on exploration, making it ideal for learning how different inference strategies impact:

โšก Latency

๐Ÿ”ฅ Throughput (tokens/sec)

๐Ÿ’พ Memory usage

๐Ÿ’ต Cost efficiency

๐Ÿค– Real-time user experience

๐ŸŽฏ What This App Teaches

Each page inside the Streamlit UI focuses on one inference concept, showing both code and performance impact:

๐Ÿงฉ 1. Batching

Demonstrates how batching multiple prompts drastically increases throughput.

Shows tokens/sec vs batch size.

๐Ÿ’ก 2. KV-Cache

Visualizes how reusing cached key/value tensors reduces decoding cost.

Demonstrates "streaming-like" speedup.

โšก 3. Speculative Decoding

Draft model generates N tokens โ†’ target model verifies.

Shows latency reduction %.

๐Ÿ”ฅ 4. vLLM Engine

Compares vanilla inference vs paged attention.

Great for understanding GPU memory efficiency.

๐Ÿค– 5. Chatbot Demo

Side-by-side inference comparison.

Helps visualize real-time responsiveness differences.

๐Ÿ“Š 6. Final Benchmark

A clean, unified benchmark runner measuring:

Latency

Throughput

Cost estimates

Stability across multiple runs

โ–ถ๏ธ Running Locally (Recommended for Testing)

  1. 1.Install dependencies pip install -r requirements.txt
  1. 1.Launch the Streamlit app streamlit run streamlit_app/Intro.py

You should now see the multipage UI load with all benchmark pages.