Kotgire58/InferenceBenchmark
๐ Inference Benchmarking Suite (Streamlit Edition)
A modular, educational benchmarking toolkit designed to teach and demonstrate modern LLM inference optimizations โ including batching, KV-cache reuse, speculative decoding, vLLM acceleration, and real-time chatbot comparisons โ with a clean Streamlit UI.
This project is structured for hands-on exploration, making it ideal for learning how different inference strategies impact:
โก Latency
๐ฅ Throughput (tokens/sec)
๐พ Memory usage
๐ต Cost efficiency
๐ค Real-time user experience
๐ฏ What This App Teaches
Each page inside the Streamlit UI focuses on one inference concept, showing both code and performance impact:
๐งฉ 1. Batching
Demonstrates how batching multiple prompts drastically increases throughput.
Shows tokens/sec vs batch size.
๐ก 2. KV-Cache
Visualizes how reusing cached key/value tensors reduces decoding cost.
Demonstrates "streaming-like" speedup.
โก 3. Speculative Decoding
Draft model generates N tokens โ target model verifies.
Shows latency reduction %.
๐ฅ 4. vLLM Engine
Compares vanilla inference vs paged attention.
Great for understanding GPU memory efficiency.
๐ค 5. Chatbot Demo
Side-by-side inference comparison.
Helps visualize real-time responsiveness differences.
๐ 6. Final Benchmark
A clean, unified benchmark runner measuring:
Latency
Throughput
Cost estimates
Stability across multiple runs
โถ๏ธ Running Locally (Recommended for Testing)
- Install dependencies pip install -r requirements.txt
- Launch the Streamlit app streamlit run streamlit_app/Intro.py
You should now see the multipage UI load with all benchmark pages.
