nareshns2004/GPU-Inference-Optimization-Lab
0
1---2title: GPU Inference Optimization Lab3emoji: ๐4colorFrom: blue5colorTo: blue6sdk: static7pinned: false8license: apache-2.09short_description: LLM Inference performance and GPU efficiency10---11 12# GPU Inference Optimization Lab13 14Measure how optimizations change LLM inference speed and GPU efficiency: precision (fp32/fp16/bf16), batch size, `torch.compile`, and later quantization and attention kernels.15 16- `lab/measure.py` โ times generation, records tokens/s and peak GPU memory17- `lab/run_experiments.py` โ runs a grid of settings and prints a comparison table18 19## Run locally20 21```bash22pip install -r requirements.txt23python lab/run_experiments.py24```25 